|
Pattern Recognition (Prof. Chanho Eom, 1 paper)
관리자 │ 2026-04-20 HIT 671 |
|---|
|
We are delighted to announce that one paper from the Perceuptual AI Lab (PAI Lab, Prof. Chanho Eom) has been accepted to Pattern Recognition. Title: Generalizable Large Language Model Based Human Keypoint Localization for Emotion Recognition Authors: Jianing Li, Xiaobin Liu, Chanho Eom, Shuang Yang, Jianzhong He, Hantao Yao, Jing Yuan Abstract: Emotion recognition aims to read face, posture and gesture, which critically requires flexible keypoint localization. However, conventional keypoint detectors only recognize predefined keypoint types, blind to unseen ones. This issue is particularly critical for affective computing where emotional states often manifest through subtle and context-dependent keypoint variations. Multi-modal Large Language Model (MLLM) has the potential to recognize unseen types of keypoints. However, existing MLLM-based methods typically predict keypoint coordinates via string generation or coordinate regression, suffering from the textual-spatial misalignment and the difficulty of long-range regression, respectively. These inefficient formulations limit the localization accuracy and efficiency. To address these issues, this paper presents a simple yet effective method, named Large Language-guided Localization Model (L3M). Different from existing formulations, L3M represents each keypoint coordinate with a single Spatial Coordinate Token (SCT) in MLLM’s vocabulary, avoiding the overhead of textual encoding and alleviating the difficulty of long-range regression. To further enhance the localization ability, we propose an visual token enhancement strategy that embeds fine-grained details into coarse-grained image tokens, enabling the model to perceive subtle cues while incurring minimal computational cost. Diverse keypoint textual instructions are additionally constructed to guide the learning of L3M. Experiments on widely used emotion recognition and keypoint localization benchmarks demonstrate the promising performance of L3M across various settings. Those results highlight that L3M not only advances keypoint localization, but also offers a more adaptive foundation for affective computing tasks to recognize diverse and fine-grained emotional cues. |
| 이전글 | ACL 2026 (Prof. YoungBin Kim, 4 papers) |
|---|---|
| 다음글 | IEEE Transactions on Image Processing (Prof. Chanho Eom, 1 paper) |