000 03552nam a2200373 a 4500
005 20260817161322.0
008 260218s20259999th mm 000 eng d
035 _a.b12476699
099 9 _aAIT Thesis no.DSAI-25-05
100 1 _aDoula, Md Shafi Ud
245 1 0 _aA multi-modal framework for context-aware plant disease classification and segmentation integrating visual and textual features
260 _aPathum Thani, Thailand :
_bAsian Institute of Technology,
_c2025
300 _a74 leaves :
_bill.+
_e1 online resource
490 1 _aThesis ;
_vno. DSAI-25-05
500 _aA thesis submitted in partial fulfillment of the requirements for the degree of Master of Engineering in Data Science and Artificial Intelligence, School of Engineering and Technology
502 _aThesis (M. Eng.) - Asian Institute of Technology, 2025
520 _aPlant diseases substantially challenge agricultural productivity and global food security. Hence, better intelligent and interpretable diagnostic frameworks are needed. An auto mated disease identification system can reduce the human effort in checking large farms, and early detection and identification will minimize the loss, which ultimately positively affects the economy. Traditional image-based deep learning models, particularly Convo lutional Neural Networks (CNNs), often struggle to distinguish visually similar diseases due to the absence of contextual information. To address these limitations, we present an innovative multi-modal deep learning framework that effectively combines visual and textual data to improve plant disease classification and segmentation. Initially, the framework incorporates a linguistically enriched Text Encoder, where disease-related descriptions are preprocessed using natural language processing (NLP) techniques to extract salient noun, numerical, adjective, and adverbial features. These refined textual representations are then encoded using a fine-tuned transformer-based language model, capturing domain-specific semantics crucial for disease differentiation. Concurrently, CNN-based Vision Encoder extract discriminative hierarchical features, which are dy namically fused with textual representations via a multi-head attention mechanism, en suring adaptive cross-modal feature alignment. Unlike conventional fusion techniques, our approach learns complex inter-dependencies between textual cues and visual pat terns, enhancing classification accuracy and segmentation precision. Finally, we demon strate our proposed framework{u2019}seffectiveness byevaluating it on the Plant Disease Diag nosis Multimodal Dataset (PDDM) and achieving state-of-the-art (SOTA) segmentation and classification performance.
650 0 _aPlant pathology
_xData processing
650 0 _aNatural language processing (Computer science)
650 0 _aAgriculture
700 0 _aChutiporn Anutariya,
_eChairperson
700 0 _aMongkol Ekpanyapong,
_eExamination Committee
700 0 _aCherdsak Kingkan,
_eExamination Committee
710 2 _aAIT Scholarship,
_eScholarship Donor
810 2 _aAsian Institute of Technology.
_tThesis ;
_vno. DSAI-25-05
856 4 0 _3Full-Text
_uhttp://203.159.5.9/ait-thesis/detail.php?q=B23568
907 _a.b12476699
_bmnarc
_ca
902 _a260309
998 _b0
_c260218
_dm
_eh
_fa
_g0
945 _lmnarc
942 _c67
909 _aBarcode : -
_bCREATED : 2026-02-18
_cRECORD # : i13573652
_dLPATRON : 0
_eLCHKIN : -
_f# RENEWALS : 0
_g# OVERDUE : 0
_hIUSE3 : 0
_iTOT CHKOUT : 0
_jTOT RENEW : 0
999 _c835
_d835