Total: 1
Audio deep reasoning demands expert-level perception and multi-step reasoning. Vanilla multi-agent paradigms struggle here due to three challenges: (1) underutilization of the complementary abilities of diverse Large Audio-Language Models (LALMs); (2) neglect of role differentiation and capability bias; and (3) sampling instability inherent to LALMs. To address these, we propose AsymAudio, a novel Capability-Aware Multi-Agent Audio Reasoning framework with three key components: (1) an Inter-Agent Collaborative Interaction Module that fuses complementary acoustic strengths across LALMs; (2) a Role-Adaptive Asymmetric Strategy that assigns hierarchical roles and tailored feedback to reduce synergistic hallucination from the bucket effect and textual compensation; and (3) an Intra-Agent Consistency Refinement Module that uses resampling and self-correction to enhance stability. AsymAudio ranked 3rd in the Interspeech 2026 Audio Deep Reasoning Challenge (Agent Track).