Total: 1
Auditory Large Language Models (ALLMs) show strong performance in audio understanding and reasoning, but their reliability is limited by hallucinations. Existing evaluation methods cast hallucination detection as binary classification, failing to capture the nuanced patterns in generative audio tasks, while mitigation methods often rely on costly fine-tuning. To address this, we propose a plug-and-play Noise-Aware In-Context Learning (NAICL) method. NAICL constructs a noise prior library, retrieves noise examples relevant to the input audio, and incorporates them as contextual priors to reduce speculative associations and generate more conservatively when acoustic evidence is insufficient. We also establish the Clotho-1K hallucination benchmark, define four types of auditory hallucinations, and introduce fine-grained evaluation metrics. Experiments show that evaluated ALLMs share similar hallucination behaviors, and NAICL reduces the hallucination rate from 26.53% to 16.98%.