Total: 1
Large Audio Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information during multi-modal fusion. This bias limits the utilization of acoustic cues and degrades audio reasoning performance. To mitigate this, we propose MATA, a novel training-free method that dynamically pushes LALMs to pay More Attention To Audio tokens within the self-attention mechanism. Specifically, MATA intervenes after raw attention scoring, targeting only the last token in intermediate layers without adding parameters or computational overhead. Experiments on MMAU and MMAR benchmarks confirm consistent performance gains. Furthermore, integrating MATA with the Qwen3-Omni-Thinking model secured second place in the Single Model Track of the Interspeech 2026 Audio Reasoning Challenge. As the only training-free approach among top solutions, MATA offers a highly efficient strategy to mitigate attention bias and advance LALM reasoning.