대학원소개

논문성과

ECCV 2026 (Prof. Jin-Hwi Park, 3 papers)

관리자 │ 2026-07-16

HIT

256

We are delighted to announce that 3 papers from CSILab (Prof. Jin-Hwi Park) have been accepted to ECCV 2026.


Title: 

Disentangling Rotation and Translation from SE(3)-Equivariant Features for Shape Assembly


Authors:

Heejun Jung, Uigeun Ahn, Jinhwi Park†, Kangil Kim†


Abstract:

3D shape assembly requires predicting 3D pose by estimating rotation and translation separately to align objects or fractures. For 3D pose prediction, prior works utilize the entangled features that capture 3D pose information well, but this mixed representation impedes generalization. To address this problem, we propose the SOT encoder that disentangles 3D pose into an SO(3)-equivariant (rotation) feature and a T(3)-equivariant (translation) feature. The SOT encoder consists of two branches: 1) a rotation branch that suppresses translation by projecting features onto the translation null space, and 2) a translation branch that suppresses rotation by translation loss and directly subtracting the SO(3)-equivariant representation from its SE(3)-equivariant counterpart. Across a variety of assembly models and diverse datasets, the SOT encoder consistently demonstrates improved generalization performance in 3D shape assembly. Furthermore, in-depth analyses indicate that the disentangled features show improved equivariant behavior with respect to the target factor while exhibiting invariant behavior with respect to the other. These results indicate that explicitly encouraging factor disentanglement is a straightforward and effective approach for 3D shape assembly.


Title: 

When Sinks Help or Hurt: Unified Framework for Attention Sink in MLLMs


Authors:

Jiho Choi, Jaemin Kim, Sanghwan Kim†, Seunghoon Hong†, Jinhwi Park†


Abstract:

Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Vision-Language Models (LVLM) remains largely unexplored: are they redundant artifacts or essential global priors? This paper first categorizes visual sinks into two distinct categories: ViT-emerged sinks (V-sinks), which propagate from the vision encoder, and LLM-emerged sinks (L-sinks), which arise within deep LLM layers. Based on the new definition, our analysis reveals a fundamental performance trade-off: while sinks effectively encode global scene-level priors, their dominance can suppress the fine-grained visual evidence required for local perception. Furthermore, we identify specific functional layers where modulating these sinks most significantly impacts downstream performance. To leverage these insights, we propose Layer-wise Sink Gating (LSG), a lightweight, plug-and-play module that dynamically scales the attention contributions of V-sink and the rest visual tokens. LSG is trained via standard next-token prediction, requiring no task-specific supervision while keeping the LVLM backbone frozen. In most layers, LSG yields improvements on representative multimodal benchmarks, effectively balancing global reasoning and precise local evidence.



Title: 

PIAvatar: Physically Interactive Avatars via Deformation Gradient Decoupling


Authors:

Sanghun Han, Mingyu Park, Jisu Shin, Seunghyun Shin, Jinhwi Park, Haegon Jeon†


Abstract:

3D human avatars have shown impressive visual fidelity driven by pose-conditioned models, yet they still lack the physical ability required for interactions with each other and environments. Although recent studies have made various attempts to incorporate physical characteristics into 3D avatars, they only exhibit limited physical deformations, often leading to constrained interaction behaviors. To resolve this issue, we present PIAvatar, a framework to simultaneously enable physically aware interactions between avatar-avatar and avatar-environment, and a non-rigid deformable human body simulation. In this work, our key insight is to decouple kinematic velocity from deformation gradient. When external forces act on avatars, the kinematic velocity induces stress which hinders the avatar's ability to achieve a desired pose. In addition, we integrate a skeletal framework within the avatar. It allows estimating its poses and real-time tracking in a closed form, even during non-rigid physical interactions. Our approach is implemented within a conventional Material Point Method framework to ensure physically consistent dynamics. We lastly evaluate the method on both human-object and human-human interaction scenarios to assess its behavior under diverse interaction settings.


이전글 Image and Vision Computing (Prof. Jinwan Park, 1 paper)
다음글 ECCV 2026 (Prof. Hyeokjun Kweon, 1 paper)