he26e@interspeech_2026@ISCA

Total: 1

#1 Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Xiang He, Chenxing Li, Jinting Wang, Yan Rong, Tianxin Xie, Zeyu Xie, Wenfu Wang, Li Liu, Dong Yu

Large Audio-Language Models (LALMs) excel at perception but lack grounded reasoning. Existing methods rely on supervised chain-of-thought (CoT) data or coarse Reinforcement Learning (RL) rewards that do not directly evaluate reasoning quality, yielding chains logically ungrounded in audio. To bridge this gap, we propose Audio-DeepThinker with two ideas. First, a hybrid reasoning similarity reward combining an LLM evaluator assessing logical path alignment and key step coverage with embedding similarity enforcing semantic alignment with reference chains. Second, a progressive two-stage curriculum enabling CoT to emerge via pure RL exploration (RL-Zero). Stage 1 uses the hybrid reward on foundational audio QA, while Stage 2 shifts to boundary cases with an LLM-only reward for reasoning diversity. Audio-DeepThinker achieves state-of-the-art on MMAR (74.0%) and MMAU-Test-Mini (78.5%), winning 1st Place in the Interspeech 2026 Audio Reasoning Challenge (Single Model Track).

Subject: INTERSPEECH.2026 - Modelling and Learning