Transformer-based aggregation for sequential face recognition (Record no. 122420)

MARC details
000 -LEADER
fixed length control field 05166nas a2200397 a 4500
005 - DATE AND TIME OF LATEST TRANSACTION
control field 20260819085037.0
008 - FIXED-LENGTH DATA ELEMENTS--GENERAL INFORMATION
fixed length control field 260219s20259999th u ms t 000 eng d
035 ## - SYSTEM CONTROL NUMBER
System control number .b12478088
099 #9 - LOCAL FREE-TEXT CALL NUMBER (OCLC)
Classification number AIT Diss. no.CS-25-01
100 1# - MAIN ENTRY--PERSONAL NAME
Personal name Jednipat Moonrinta
245 10 - TITLE STATEMENT
Title Transformer-based aggregation for sequential face recognition
260 ## - PUBLICATION, DISTRIBUTION, ETC.
Place of publication, distribution, etc. Pathum Thani, Thailand :
Name of publisher, distributor, etc. Asian Institute of Technology,
Date of publication, distribution, etc. 2025
300 ## - PHYSICAL DESCRIPTION
Extent 152 leaves :
Other physical details ill. +
Accompanying material 1 online resource
490 1# - SERIES STATEMENT
Series statement Dissertation ;
Volume/sequential designation no. CS-25-01
500 ## - GENERAL NOTE
General note A dissertation submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy in Computer Science
502 ## - DISSERTATION NOTE
Dissertation note Thesis (Ph.D.) - Asian Institute of Technology, 2025
520 ## - SUMMARY, ETC.
Summary, etc. In the past decade, advancements in face recognition have achieved accuracy levels com parable to human performance. The most effective models perform face embedding, in which a cropped face image is projected into a vector space in a way that embeddings of the same individual{u2019}s face under different imaging conditions tend to lie close together, and embeddings of different individuals{u2019} faces tend to lie far apart. Face embedding models developed using machine learning have driven success in various real-world applications such as face verification for authentication. However, these systems typically require high-quality input images for both gallery and probe samples, an assumption that does not always hold in real-world scenarios. Variations in image quality, illumination, face pose, and occlusions can significantly affect performance. To address the challenge of noisy face embeddings caused by inconsistent image quality, we propose a deep learning-based approach that aggregates multiple face embeddings, either sequential or non-sequential, to construct a more robust face representation for verification.First, we develop a Transformer-based model that processes face embeddings extracted from a high-performance face feature extractor. The model takes a batch of embeddings as input and generates a refined face representation for verification.Second, we introduce an adaptive triplet loss function that extends the traditional triplet loss function for face embedding models by applying separate margin thresholds for positive and negative pairs. We evaluate both fixed margins and dynamically adjusted margins that adapt to the data on the fly. The adaptive triplet loss mitigates excessive penalization of differences between positive pairs and similarities between negative pairs.Third, we implement a two-stage training strategy comprising pre-training and fine tuning phases. In the pre-training phase, the model is trained using root mean squared loss to produce embeddings similar to those obtained via average pooling. In the fine tuning phase, the above-mentioned adaptive triplet loss is employed to further refine the model. In a series of experiments, we train our proposed model using this two-stage training approach on the YouTubeFaces dataset and evaluate its performance on multiple benchmark datasets, including IMFDB, CASIA-WebFace, CCVID, IJB-B, and IJB-C.The experimental results demonstrate that our method outperforms the baseline on most datasets. Our model generalizes well to unseen datasets and performs exceptionally well on the IMFDB dataset, where significant variations exist in face pose, lighting condi tions, and scene contexts within a single identity.Furthermore, to assess the model{u2019}s effectiveness in real-world deployment, we integrate our Transformer-based aggregation model into an elder care scenario using a telehealth dataset. In this setting, the model must verify individuals using video sequences captured in uncontrolled environments, with natural lighting, and high face pose variations. Our method shows improvement over the average pooling baseline. This confirms the practical utility of our approach for improving face verification reliability in challenging real-world conditions. We conclude that the methodology is an effective way to exploit the availability of multiple images of an individual when performing face verification under adverse conditions.
650 #0 - SUBJECT ADDED ENTRY--TOPICAL TERM
Topical term or geographic name entry element Human face recognition (Computer science)
650 #0 - SUBJECT ADDED ENTRY--TOPICAL TERM
Topical term or geographic name entry element Human-computer interaction
650 #0 - SUBJECT ADDED ENTRY--TOPICAL TERM
Topical term or geographic name entry element Computer vision
700 1# - ADDED ENTRY--PERSONAL NAME
Personal name Dailey, Mathew N.,
Relator term Chairperson
700 0# - ADDED ENTRY--PERSONAL NAME
Personal name Huynh, Trung Luong,
Relator term Examination Committee
700 0# - ADDED ENTRY--PERSONAL NAME
Personal name Chaklam Silpasuwanchai,
Relator term Examination Committee
700 0# - ADDED ENTRY--PERSONAL NAME
Personal name Adisorn Lertsinsrubtavee,
Relator term Examination Committee
710 2# - ADDED ENTRY--CORPORATE NAME
Corporate name or jurisdiction name as entry element Royal Thai Government,
Relator term Scholarship Donor
710 2# - ADDED ENTRY--CORPORATE NAME
Corporate name or jurisdiction name as entry element AIT Fellowship,
Relator term Scholarship Donor
810 2# - SERIES ADDED ENTRY--CORPORATE NAME
Corporate name or jurisdiction name as entry element Asian Institute of Technology.
Title of a work Dissertation ;
Volume/sequential designation no. CS-25-01
856 40 - ELECTRONIC LOCATION AND ACCESS
Materials specified Full-Text
Uniform Resource Identifier <a href="http://203.159.5.9/ait-thesis/Viewer/viewer.php?id=B23605">http://203.159.5.9/ait-thesis/Viewer/viewer.php?id=B23605</a>
907 ## - LOCAL DATA ELEMENT G, LDG (RLIN)
a .b12478088
b mnarc
c a
902 ## - LOCAL DATA ELEMENT B, LDB (RLIN)
a 260227
998 ## - LOCAL CONTROL INFORMATION (RLIN)
Operator's initials, OID (RLIN) 0
Cataloger's initials, CIN (RLIN) 260226
First date, FD (RLIN) m
-- h
-- a
-- 0
945 ## - LOCAL PROCESSING INFORMATION (OCLC)
l mnarc
942 ## - ADDED ENTRY ELEMENTS (KOHA)
Koha item type 67-Electronic Resource
909 ## - LOCAL ITEMS USED
Barcode Barcode : -
CREATED CREATED : 2026-02-19
RECORD Id RECORD # : i13575016
LPATRON LPATRON : 0
LCHKIN LCHKIN : -
RENEWALS # RENEWALS : 0
-- # OVERDUE : 0
-- IUSE3 : 0
-- TOT CHKOUT : 0
-- TOT RENEW : 0
Holdings
Withdrawn status Lost status Damaged status Not for loan Home library Current library Shelving location Date acquired Total checkouts Full call number Date last seen Copy number Price effective from Koha item type
      Available for Loans Asian Institute of Technology Library Asian Institute of Technology Library Archives 19/08/2026   AIT Diss. no.CS-25-01 19/08/2026 1 19/08/2026 67-Electronic Resource
คัดลอกแล้ว!