000 03861nas a2200469 a 4500
005 20260817173419.0
008 240229s2022 th u m tt 000 a eng d
035 _a.b12421662
099 _aAIT Thesis no.DSAI-22-04
100 0 _aChanapa Pananookooln
245 1 0 _aComparing selective masking methods for depression detection in social media
260 _aPathum Thani, Thailand :
_bAsian Institute of Technology,
_c2022
300 _a47 leaves :
_bill.
490 1 _aThesis ;
_vno. DSAI-22-04
500 _aA thesis submitted in partial fulfillment of the requirements for the degree of Master of Science in Data Science and Artificial Intelligence, School of Engineering and Technology
502 _aThesis (M. Sc.) - Asian Institute of Technology, 2022
520 _aIdentifying those at risk for depression is a crucial issue in which social media provides an excellent platform for examining the linguistic patterns of depressed individuals. A significant challenge in a depression classification problem is ensuring that the predic tion model is not overly dependent on keywords, such that it fails to predict when key words are unavailable. One promising approach is masking, i.e., by masking important words selectively and asking the model to predict the masked words, the model is forced to learn the context rather than the keywords. This study evaluates seven masking tech niques, such as random masking, log-odds ratio, and the use of attention scores. In ad dition, whether to predict the masked words during pretraining or fine-tuning phase was also examined. Last, six class imbalance ratios were compared to determine the robust ness of the masked selection methods. Key findings demonstrated that selective masking generally outperforms random masking in terms of classification accuracy. In addition, the most accurate and robust models were identified. Our research also indicated that re constructing the masked words during the pre-training phase is more advantageous than during the fine-tuning phase. Further discussion and implications were made. This is the first study to comprehensively compare masking selection methods, which has broad implications for the field of depression classification and the general NLP. Our code can be found in: https://github.com/chanapapan/Depression-Detection
650 0 _aSocial media
_xData processing
650 0 _aMachine learning
650 0 _aNeural networks (Computer science)
700 0 _aChaklam Silpasuwanchai,
_eChairperson
700 1 _aDailey, Matthew N.,
_eExamination Committee
700 0 _aMongkol Ekpanyapong,
_eExamination Committee
710 2 _aHis Majesty the King{u2019}s Scholarships (Thailand),
_eScholarship Donor
810 2 _aAsian Institute of Technology.
_tThesis ;
_vno. DSAI-22-04
856 4 0 _3Full-Text
_uhttp://203.159.5.9/ait-thesis/Viewer/viewer.php?id=B20417
907 _a.b12421662
_bmnait
_ca
902 _a250428
998 _b0
_c240313
_dm
_ea
_fa
_g0
945 _lmnarc
945 _lmnarc
945 _lmnait
945 _lmnait
942 _c67
942 _c40
942 _c20
909 _aBarcode : -
_bCREATED : 2024-02-29
_cRECORD # : i13487875
_dLPATRON : 0
_eLCHKIN : -
_f# RENEWALS : 0
_g# OVERDUE : 0
_hIUSE3 : 0
_iTOT CHKOUT : 0
_jTOT RENEW : 0
909 _aBarcode : 30050120899066
_bCREATED : 2025-07-03
_cRECORD # : i13537118
_dLPATRON : 0
_eLCHKIN : -
_f# RENEWALS : 0
_g# OVERDUE : 0
_hIUSE3 : 0
_iTOT CHKOUT : 0
_jTOT RENEW : 0
909 _aBarcode : 30050120425151
_bCREATED : 2025-08-04
_cRECORD # : i1354293x
_dLPATRON : 0
_eLCHKIN : -
_f# RENEWALS : 0
_g# OVERDUE : 0
_hIUSE3 : 0
_iTOT CHKOUT : 0
_jTOT RENEW : 0
909 _aBarcode : 30050120425144
_bCREATED : 2025-08-04
_cRECORD # : i13542941
_dLPATRON : 0
_eLCHKIN : -
_f# RENEWALS : 0
_g# OVERDUE : 0
_hIUSE3 : 0
_iTOT CHKOUT : 0
_jTOT RENEW : 0
999 _c19836
_d19836