Total: 1
Large Audio Language Models (LALMs) suffer from Neutral Bias, failing to capture subtle emotional nuances. Existing attempts to address this through text augmentation usually introduce acoustic hallucinations. To address these issues without costly human annotation, we propose a data-centric paradigm: Semantic Drift and Discriminative Re-ranking. Our method employs an LLM to perform multi-dimensional semantic drift, exploring diverse emotional hypotheses beyond neutral baselines. Subsequently, a discriminative judge model filters drifted hypotheses via hard-negative contrastive learning, grounding them in raw acoustic signals. Experiments confirm this hypothesize-and-verify mechanism transforms coarse, neutral-biased labels into precise, acoustically grounded descriptions, enriching fine-grained emotional detail and downstream expressiveness.