| 000 | 03146nas a2200397 a 4500 | ||
|---|---|---|---|
| 005 | 20260818095404.0 | ||
| 008 | 240223s2023 th u m tt 000 a eng d | ||
| 035 | _a.b12420591 | ||
| 099 | _aAIT Thesis no.CS-23-05 | ||
| 100 | 1 | _aLamsal, Diwas | |
| 245 | 1 | 0 |
_aEfficient posevit : _bexploring efficient vision transformers and training strategies for fast human pose estimation |
| 260 |
_aPathum Thani, Thailand : _bAsian Institute of Technology, _c2023 |
||
| 300 |
_a74 leaves : _bill. |
||
| 490 | 1 |
_aThesis ; _vno.CS-23-05 |
|
| 500 | _aA thesis submitted in partial fulfillment of the requirements for the degree of Master of Engineering in Computer Science | ||
| 502 | _aThesis (M. Eng.) - Asian Institute of Technology, 2023 | ||
| 520 | _aHuman pose estimation (HPE) serves as a critical component in various downstream applications, including activity recognition and human-computer interaction. However, it faces challenges such as occlusion, complex body poses, diverse clothing and environments, and low-quality images. State-of-the-art HPE methods adept at addressing these challenges often prove to be computationally expensive to run on low-cost edge devices. On the other hand, the efficient models designed to run on edge devices are not as accurate, especially when faced with occlusion. In this thesis, I explored the use of computationally efficient vision transformer (ViT) backbones for the task of HPE to produce an efficient model capable of running on edge devices while also addressing common HPE challenges. Specifically, I replaced the ViT encoder in ViTPose with effi cient ViT variants, and performed a comprehensive comparison. The findings indicate that the EdgeNeXt-Base model offers an optimal speed-accuracy trade-off among the evaluated models. Through knowledge distillation and multi-task training strategies, the resulting models exhibit improved performance, particularly in handling occlusion. Additionally, I produced a wheelchair-based human pose estimation and activity recognition dataset aimed at developing a system that effectively monitors the activities of people sitting in wheelchairs. | ||
| 650 | 0 | _aImage analysis | |
| 650 | 0 | _aHuman-computer interaction | |
| 700 | 1 |
_aDailey, Matthew N., _eChairperson |
|
| 700 | 0 |
_aMongkol Ekpanyapong, _eExamination Committee |
|
| 700 | 0 |
_aChaklam Silpasuwanchai, _eExamination Committee |
|
| 710 | 2 |
_aAIT Scholarships, _eScholarship Donor |
|
| 810 | 2 |
_aAsian Institute of Technology. _tThesis : _vno.CS-23-05 |
|
| 856 | 4 | 0 |
_3Full-Text _uhttp://203.159.5.9/ait-thesis/Viewer/viewer.php?id=B20405 |
| 907 |
_a.b12420591 _bmnait _ca |
||
| 902 | _a250303 | ||
| 998 |
_b0 _c240312 _dm _ea _fa _g0 |
||
| 945 | _lmnarc | ||
| 945 | _lmnarc | ||
| 942 | _c67 | ||
| 942 | _c40 | ||
| 909 |
_aBarcode : - _bCREATED : 2024-02-23 _cRECORD # : i13486640 _dLPATRON : 0 _eLCHKIN : - _f# RENEWALS : 0 _g# OVERDUE : 0 _hIUSE3 : 0 _iTOT CHKOUT : 0 _jTOT RENEW : 0 |
||
| 909 |
_aBarcode : 30050120900476 _bCREATED : 2025-03-03 _cRECORD # : i13535183 _dLPATRON : 0 _eLCHKIN : - _f# RENEWALS : 0 _g# OVERDUE : 0 _hIUSE3 : 0 _iTOT CHKOUT : 0 _jTOT RENEW : 0 |
||
| 999 |
_c35568 _d35568 |
||