Leveraging phonological clustering for word-level Bangla Sign language recognition

By: Call Number: AIT Thesis no.CS-25-01 Contributor(s): Material type: SeriesSeries: Asian Institute of Technology. Thesis ; no.CS-25-01Publication details: Pathum Thani, Thailand : Asian Institute of Technology, 2025Description: 63 leaves : ill.+ 1 online resourceSubject(s): Online resources: Dissertation note: Thesis (M. Sc.) - Asian Institute of Technology, 2025 Summary: Bangla Sign Language (BdSL) recognition presents multifaceted challenges due to signer diversity and spatiotemporal variability. While flat classification pipelines are widely used, they often overlook the underlying phonological relationships among signs that can inform more structured and accurate recognition. To address this gap, we propose a hierarchical recognition framework that integrates phonological clustering into a Bidirectional Long Short-Term Memory (Bi-LSTM)-based sequence modeling pipeline. First, baseline classification{u2014}referring to a flat, non-clustered recognition approach{u2014}is performed us ing five Bi-LSTM configurations of increasing complexity to assess the trade-off between accuracy and model size. The resulting confusion matrices are analyzed to identify sign pairs with high misclassification rates, revealing underlying phonological similarities. Based on this analysis, a confusion-matrix driven clustering strategy is employed to group visually and phonologically similar signs. Cluster-specific feature engineering is then applied, and the same Bi-LSTM architecture is retrained separately within each cluster. Experiments are conducted on a curated 50-class subset of the SignBD-Word dataset. In the baseline setting, the most complex model (Bi-LSTM-1) achieves 91.10% accuracy with 2.6 million parameters. With the proposed confusion-matrix driven clustering architecture, four out of six clusters employing the lightweight Bi-LSTM classifiers outperform the baseline model, reaching up to 94.58% accuracy. Remarkably, the lightweight Bi-LSTM-4 model{u2014}with only 373K parameters (14.4% of Bi LSTM-1){u2014}achieves a weighted average accuracy of 92.60% on cluster-level classification, surpassing the baseline by1.5percentagepoints. Thisreflectsan85.6%reduction in a number of parameters, demonstrating that efficient cluster-specific feature engineering, which allows each model to capture nuanced patterns within each cluster can improve overall predictive accuracy.
Tags from this library: No tags from this library for this title. Log in to add tags.
Star ratings
    Average rating: 0.0 (0 votes)
Holdings
Cover image Item type Current library Home library Collection Shelving location Call number Materials specified Vol info URL Copy number Status Notes Date due Barcode Item holds Item hold queue priority Course reserves
67-Electronic Resource Asian Institute of Technology Library Archives AIT Thesis no.CS-25-01 (Browse shelf(Opens below)) 1 Not for loan

A thesis submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science

Thesis (M. Sc.) - Asian Institute of Technology, 2025

Bangla Sign Language (BdSL) recognition presents multifaceted challenges due to signer diversity and spatiotemporal variability. While flat classification pipelines are widely used, they often overlook the underlying phonological relationships among signs that can inform more structured and accurate recognition. To address this gap, we propose a hierarchical recognition framework that integrates phonological clustering into a Bidirectional Long Short-Term Memory (Bi-LSTM)-based sequence modeling pipeline. First, baseline classification{u2014}referring to a flat, non-clustered recognition approach{u2014}is performed us ing five Bi-LSTM configurations of increasing complexity to assess the trade-off between accuracy and model size. The resulting confusion matrices are analyzed to identify sign pairs with high misclassification rates, revealing underlying phonological similarities. Based on this analysis, a confusion-matrix driven clustering strategy is employed to group visually and phonologically similar signs. Cluster-specific feature engineering is then applied, and the same Bi-LSTM architecture is retrained separately within each cluster. Experiments are conducted on a curated 50-class subset of the SignBD-Word dataset. In the baseline setting, the most complex model (Bi-LSTM-1) achieves 91.10% accuracy with 2.6 million parameters. With the proposed confusion-matrix driven clustering architecture, four out of six clusters employing the lightweight Bi-LSTM classifiers outperform the baseline model, reaching up to 94.58% accuracy. Remarkably, the lightweight Bi-LSTM-4 model{u2014}with only 373K parameters (14.4% of Bi LSTM-1){u2014}achieves a weighted average accuracy of 92.60% on cluster-level classification, surpassing the baseline by1.5percentagepoints. Thisreflectsan85.6%reduction in a number of parameters, demonstrating that efficient cluster-specific feature engineering, which allows each model to capture nuanced patterns within each cluster can improve overall predictive accuracy.

There are no comments on this title.

to post a comment.
คัดลอกแล้ว!