SADGR: Adaptive Cross-modal Emotion Recognition via Self-supervised Alignment and Dynamic Gating

  • Zhang, Junjun
  • Mao, Jianjing
  • Hou, Yanyang
  • Noh, Giseop
  • Zhang, Fengxi
Citations

SCOPUS

0

초록

In this paper, we propose SADGR, an adaptive cross-modal emotion recognition framework that follows an 'align first, modulate second, and fuse later' paradigm. First, we introduce a self-supervised text-audio contrastive alignment (STA-CA) stage, which leverages large-scale unlabeled audio-text data to map heterogeneous modalities into a shared semantic space prior to fusion. Second, we design an audio-guided visual gating (AG-VR) mechanism, where emotionally salient acoustic cues dynamically modulate visual representations, suppressing redundant information while enhancing emotion-relevant patterns. Extensive experiments on the CMU-MOSI and CMU-MOSEI benchmarks demonstrate that SADGR consistently outperforms state-of-the-art methods across both regression and classification tasks, achieving superior performance in ACC-7, ACC-2, and F1-score while maintaining competitive correlation results. These results indicate that explicit self-supervised alignment combined with dynamic cross-modal gating provides an effective and practical solution for robust multimodal emotion recognition. © 2013 IEEE.

키워드

Contrastive learningcross-modal emotion analysisgating mechanism
제목
SADGR: Adaptive Cross-modal Emotion Recognition via Self-supervised Alignment and Dynamic Gating
저자
Zhang, JunjunMao, JianjingHou, YanyangNoh, GiseopZhang, Fengxi
DOI
10.1109/ACCESS.2026.3663574
발행일
2026
유형
Article in press
저널명
IEEE Access
14
페이지
28229 ~ 28244