Total: 1
Large Audio-Language Models (LALMs) excel at perception but lack grounded reasoning. Existing methods rely on supervised chain-of-thought (CoT) data or coarse Reinforcement Learning (RL) rewards that do not directly evaluate reasoning quality, yielding chains logically ungrounded in audio. To bridge this gap, we propose Audio-DeepThinker with two ideas. First, a hybrid reasoning similarity reward combining an LLM evaluator assessing logical path alignment and key step coverage with embedding similarity enforcing semantic alignment with reference chains. Second, a progressive two-stage curriculum enabling CoT to emerge via pure RL exploration (RL-Zero). Stage 1 uses the hybrid reward on foundational audio QA, while Stage 2 shifts to boundary cases with an LLM-only reward for reasoning diversity. Audio-DeepThinker achieves state-of-the-art on MMAR (74.0%) and MMAU-Test-Mini (78.5%), winning 1st Place in the Interspeech 2026 Audio Reasoning Challenge (Single Model Track).