zhang26ca@interspeech_2026@ISCA

Total: 1

#1 ACR-Net: Mitigating Semantic Dominance via Contrastive Acoustic-Semantic Decoupling [PDF] [Copy] [Kimi] [REL]

Authors: Mengke Zhang, Yanda Shao, Tianhe Wu, Kai Feng

Speech Emotion Recognition (SER) models and Audio LLMs fail when vocal tone contradicts textual semantics. Extensive training on congruent data engenders a rigid semantic prior, causing "Semantic Dominance", a severe text bias during cross-modal conflicts. We introduce the ASPIRE benchmark of adversarial audio-text pairs across four conflict types, alongside two metrics: Semantic Overconfidence Penalty (SOP) and Latent Decoupling Degree (LDD). Unlike fusion methods that merely re-weight collapsed features, our Acoustic Conflict Resolution Network (ACR-Net) uses Cross-Modal Attention and Contrastive Decoupling Loss to disentangle contradictory representations into orthogonal latent spaces. Experiments show existing models fail under polarity conflicts due to extreme semantic overconfidence. Conversely, ACR-Net preserves acoustic fidelity, maintaining high LDD to drive SOP near zero, achieving superior Acoustic Accuracy (ACC) in both adversarial and standard scenarios.

Subject: INTERSPEECH.2026 - Speech Recognition