Comparing selective masking methods for depression detection in social media

By: Call Number: AIT Thesis no.DSAI-22-04 Contributor(s): Material type: SeriesSeries: Asian Institute of Technology. Thesis ; no. DSAI-22-04Publication details: Pathum Thani, Thailand : Asian Institute of Technology, 2022Description: 47 leaves : illSubject(s): Online resources: Dissertation note: Thesis (M. Sc.) - Asian Institute of Technology, 2022 Summary: Identifying those at risk for depression is a crucial issue in which social media provides an excellent platform for examining the linguistic patterns of depressed individuals. A significant challenge in a depression classification problem is ensuring that the predic tion model is not overly dependent on keywords, such that it fails to predict when key words are unavailable. One promising approach is masking, i.e., by masking important words selectively and asking the model to predict the masked words, the model is forced to learn the context rather than the keywords. This study evaluates seven masking tech niques, such as random masking, log-odds ratio, and the use of attention scores. In ad dition, whether to predict the masked words during pretraining or fine-tuning phase was also examined. Last, six class imbalance ratios were compared to determine the robust ness of the masked selection methods. Key findings demonstrated that selective masking generally outperforms random masking in terms of classification accuracy. In addition, the most accurate and robust models were identified. Our research also indicated that re constructing the masked words during the pre-training phase is more advantageous than during the fine-tuning phase. Further discussion and implications were made. This is the first study to comprehensively compare masking selection methods, which has broad implications for the field of depression classification and the general NLP. Our code can be found in: https://github.com/chanapapan/Depression-Detection
Tags from this library: No tags from this library for this title. Log in to add tags.
Star ratings
    Average rating: 0.0 (0 votes)
Holdings
Cover image Item type Current library Home library Collection Shelving location Call number Materials specified Vol info URL Copy number Status Notes Date due Barcode Item holds Item hold queue priority Course reserves
67-Electronic Resource Asian Institute of Technology Library Archives AIT Thesis no.DSAI-22-04 (Browse shelf(Opens below)) Not for loan
40-Archives Asian Institute of Technology Library Archives AIT Thesis no.DSAI-22-04 (Browse shelf(Opens below)) 1 Available 30050120899066
20-AIT Publication Asian Institute of Technology Library AIT Publications AIT Thesis no.DSAI-22-04 (Browse shelf(Opens below)) 1 Available 30050120425151
20-AIT Publication Asian Institute of Technology Library AIT Publications AIT Thesis no.DSAI-22-04 (Browse shelf(Opens below)) 2 Available 30050120425144

A thesis submitted in partial fulfillment of the requirements for the degree of Master of Science in Data Science and Artificial Intelligence, School of Engineering and Technology

Thesis (M. Sc.) - Asian Institute of Technology, 2022

Identifying those at risk for depression is a crucial issue in which social media provides an excellent platform for examining the linguistic patterns of depressed individuals. A significant challenge in a depression classification problem is ensuring that the predic tion model is not overly dependent on keywords, such that it fails to predict when key words are unavailable. One promising approach is masking, i.e., by masking important words selectively and asking the model to predict the masked words, the model is forced to learn the context rather than the keywords. This study evaluates seven masking tech niques, such as random masking, log-odds ratio, and the use of attention scores. In ad dition, whether to predict the masked words during pretraining or fine-tuning phase was also examined. Last, six class imbalance ratios were compared to determine the robust ness of the masked selection methods. Key findings demonstrated that selective masking generally outperforms random masking in terms of classification accuracy. In addition, the most accurate and robust models were identified. Our research also indicated that re constructing the masked words during the pre-training phase is more advantageous than during the fine-tuning phase. Further discussion and implications were made. This is the first study to comprehensively compare masking selection methods, which has broad implications for the field of depression classification and the general NLP. Our code can be found in: https://github.com/chanapapan/Depression-Detection

There are no comments on this title.

to post a comment.
คัดลอกแล้ว!