ahn26@interspeech_2026@ISCA

Total: 1

#1 Context-Adaptive Automated Audio Captioning with Symmetric Dual-MoE and Dynamic Reward Routing [PDF] [Copy] [Kimi] [REL]

Authors: Seyun Ahn, Joon-Hyuk Chang

Deep learning–based automated audio captioning (AAC) typically optimizes a uniform objective, limiting contextual adaptability and fine-grained audio–text alignment. Although reinforcement learning improves evaluation metrics, existing approaches rely on simple and static reward formulations. We propose a context-adaptive AAC framework based on a symmetric dual mixture-of-experts (MoE) architecture optimized with group relative policy optimization (GRPO). On the policy side, LoRA-based experts are integrated into a pretrained BART decoder and activated via cross-attention routing conditioned on acoustic context. On the reward side, parallel experts evaluate semantic relevance, grammatical correctness, lexical diversity, and audio–text alignment, while a context-aware router adaptively weights these signals. Experiments on Clotho and Audio-Caps show improvements in caption quality and human preference. Generated captions are available on the https://symmetric-moe.github.io/.

Subject: INTERSPEECH.2026 - Modelling and Learning