상세 보기
SADGR: Adaptive Cross-modal Emotion Recognition via Self-supervised Alignment and Dynamic Gating
- Zhang, Junjun;
- Mao, Jianjing;
- Hou, Yanyang;
- Noh, Giseop;
- Zhang, Fengxi
SCOPUS
0초록
In this paper, we propose SADGR, an adaptive cross-modal emotion recognition framework that follows an 'align first, modulate second, and fuse later' paradigm. First, we introduce a self-supervised text-audio contrastive alignment (STA-CA) stage, which leverages large-scale unlabeled audio-text data to map heterogeneous modalities into a shared semantic space prior to fusion. Second, we design an audio-guided visual gating (AG-VR) mechanism, where emotionally salient acoustic cues dynamically modulate visual representations, suppressing redundant information while enhancing emotion-relevant patterns. Extensive experiments on the CMU-MOSI and CMU-MOSEI benchmarks demonstrate that SADGR consistently outperforms state-of-the-art methods across both regression and classification tasks, achieving superior performance in ACC-7, ACC-2, and F1-score while maintaining competitive correlation results. These results indicate that explicit self-supervised alignment combined with dynamic cross-modal gating provides an effective and practical solution for robust multimodal emotion recognition. © 2013 IEEE.
키워드
- 제목
- SADGR: Adaptive Cross-modal Emotion Recognition via Self-supervised Alignment and Dynamic Gating
- 저자
- Zhang, Junjun; Mao, Jianjing; Hou, Yanyang; Noh, Giseop; Zhang, Fengxi
- 발행일
- 2026
- 유형
- Article in press
- 저널명
- IEEE Access
- 권
- 14
- 페이지
- 28229 ~ 28244