상세 보기
초록
Recent multimodal emotion recognition (MER) approaches often assume that fusing audio and video modalities through cross-attention inherently yields superior feature representations for MER. However, our analysis reveals that different emotions manifest more prominently in specific modalities, leading to frequent misclassifications when an unsuitable modality dominates. Building on this, we propose a framework that dynamically adjusts each modality's contribution using weighted integration of self & cross-attention for emotion, called WISE-Net. In addition, as certain emotions benefit from merging complementary audio-video information, we introduce a cross-label emotion matching task to explicitly capture fine-grained emotion-related distinctions across modalities. Experiments confirm that WISE-Net surpasses state-of-the-art methods. Ablation studies highlight the significance of balancing self-attention for modality expressiveness and cross-attention for emotion-aware feature fusion.
키워드
- 제목
- Leveraging dynamic feature fusion of self & cross-attention for robust multimodal emotion recognition
- 저자
- Kim, Jian; Kang, Hyunoh; Pyeon, Haneul; Jung, Dahuin
- 발행일
- 2026-04
- 유형
- Article
- 저널명
- ICT Express
- 권
- 12
- 호
- 2
- 페이지
- 306 ~ 310