대학원소개

논문성과

ICASSP 2026 (Prof. Chanho Eom, 1 paper)

관리자 │ 2026-03-16

HIT

621

We are delighted to announce that one paper from the Perceuptual AI Lab (PAI Lab, Prof. Chanho Eom) has been accepted to the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026).


Title: 

COVA: Text-Guided Composed Retrieval for Audio-Visual Content


Authors:

Gyuwon Han*, Young Kyun Jang*, Chanho Eom


Abstract:

Composed Video Retrieval (CoVR) aims to retrieve a target video from a large gallery using a reference video and a textual query specifying visual modifications. However, existing benchmarks consider only visual changes, ignoring videos that differ in audio despite visual similarity. To address this limitation, we introduce Composed retrieval for Video with its Audio (COVA), a new retrieval task that accounts for both visual and auditory variations. To support this, we construct AV-Comp, a benchmark of video pairs with cross-modal changes and textual queries describing the differences, enabling retrieval based on audio as well. We also propose AVT Compositional Fusion (AVT), which integrates video, audio, and text features by selectively aligning the query to the most relevant modality. AVT outperforms traditional unimodal fusion and serves as a strong baseline for COVA.



이전글 ICASSP 2026 (Prof. Jihun Kim, 2 papers)
다음글 ICLR 2026 (Prof. Jihyong Oh, 1 paper)