| Total: 176
Curriculum learning (CL) structures training from simple to complex samples, facilitating progressive learning. However, existing CL approaches for emotion recognition often rely on heuristic, data-driven, or model-based definitions of sample difficulty, neglecting the difficulty for human perception, a critical factor in subjective tasks like emotion recognition. We propose CHUCKLE (Crowdsourced Human Understanding Curriculum for Knowledge Led Emotion Recognition), a perception-driven CL framework that leverages annotator agreement and alignment in crowd-sourced datasets to define sample difficulty, under the assumption that clips challenging for humans are similarly hard for neural networks. Experimental results suggest that CHUCKLE enhances the performance of LSTMs and Transformers over non-curriculum baselines, while reducing the number of gradient updates, thereby enhancing both training efficiency and model robustness in both subject-dependent and subject-independent settings.
Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holistic naturalness, leaving fine-grained paralinguistic distinctions underexplored. We introduce PARAPAIRAUDIOBENCH, a pairwise benchmark of 5,175 audio pairs across five paralinguistic dimensions: Style, Rate, Emphasis, Age, and Gender. Our experiments show that current LALM judges still lag behind human judgments by 32%p on average and exhibit severe calibration failures, particularly in Tie cases where the correct decision is to abstain. To further analyze lexical versus acoustic reliance, the benchmark includes both same-transcript and cross-transcript conditions. PARAPAIRAUDIOBENCH enables multi-dimensional, calibration-aware assessment of the reliability of LALM-as-a-Judge for paralinguistic speech evaluation.
Hormonal changes throughout the menstrual cycle can affect speech production. However, studies investigating differences in acoustic parameters have found inconsistent results. This could be because changes are likely subtle and multivariate. We therefore apply handcrafted and embedding-based features, combined with machine learning methods. We use a German read speech dataset of 76 participants in two cycle phases (ovulation and luteal). Additionally, we investigate whether speaker-level accuracy is correlated with hormonal levels, the magnitude of hormonal changes between phases, or age. Our results show small, non-significant effects of loudness, formant amplitudes, H1-H2, and spectral flux. We achieve 62.5 % accuracy with handcrafted features, but chance-level performance with learnt embeddings. Speaker-level accuracies show no correlation with hormonal levels or age. We suggest using more cycle phases and personalisation techniques in future research.
Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic features and create a dataset of over 1.2M Audio-QA pairs. We develop ParA-LLM, trained with a two-stage curriculum: first on single-attribute questions to build foundational knowledge, then on multi-attribute questions for joint reasoning over speaker and acoustic characteristics. We also release ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories, where frontier models like GPT-4o-Audio achieve only 36% accuracy. ParA-LLM surpasses state-of-the-art Audio LLMs by 7.5% on ParA-Bench, with additional gains of 1.13% on MMAU-Pro Speech and 7.49% on MMAR Speech.
Text-to-image models ignore paralinguistic cues. Pitch, rate, and emotional inflection shape how listeners visualize a scene, yet audio-to-image methods treat speech as a single opaque vector. NovaDiffusion conditions diffusion-based synthesis on emotion-correlated prosodic features extracted alongside general audio representations. It comprises (1) Prosody-CLIP, trained on 168K human-annotated speech-emotion-image triplets to align prosody-enriched embeddings with visual representations; (2) a 280M-parameter distilled U-Net with multi-scale fusion; and (3) a decoupled cross-attention adapter injecting prosodic embeddings via IP-Adapter. On RAVDESS (8 emotions, chance 12.5%), NovaDiffusion achieves 71.3% ECA, outperforming SonicDiffusion by 23.1pp; on held-out IEMOCAP speakers it reaches 63.4% vs. 41.7%. ECA relies on facial expression classification and does not capture scene-level emotion. Scope is limited to emotion-correlated prosody; full prosodic control remains open.
Accentedness and comprehensibility scales are widely used to evaluate pronunciation development in second language learners. However, such assessments rely heavily on human rater evaluations. This study investigates whether a speech large language model (LLM) can approximate human judgments of accentedness and comprehensibility. We first compare correlations between LLM scores and human ratings. We then apply linear mixed-effects models to examine whether LLM scores capture learner progress across pre-and post-test conditions. Finally, by combining segmental and suprasegmental measures with Lasso regression, we analyze whether the LLM employs acoustic cues similar to those human raters rely on when assigning scores. The results show that the LLM scores are moderately correlated with human ratings, capture pre-post test progress, and exhibit overlapping features with human evaluations. However, future research could explore fine-tuning the LLM and incorporating linguistic knowledge.
Affective computing has largely followed an individual-state paradigm, extracting discrete emotion labels or arousal/valence from isolated speakers. We argue this framing is incomplete for interaction. Drawing on affective resonance and vitality-contour accounts, we propose a relational framework in which the primary unit of affective analysis is the interactional field constituted within vocal dynamics. As a proof of concept, we present a preliminary empirical study using continuous self-supervised speech representations to detect directional expressive coupling in multi-party conversation. Coupling is regime-specific, concentrated at sub-second timescales, and collapses under exclusive-speech negative controls, consistent with a relational account of affective dynamics. We introduce design frameworks for Artificial Affective Resonance Intelligence grounded in Affective Resonance Dynamic Ontologies, supported by null-calibrated directional coupling analyses across interaction regimes.
Unlike Speech Emotion Recognition (SER), which predicts utterance-level labels, Speech Emotion Diarization (SED) models fine-grained temporal dynamics. However, learning reliable emotion boundaries from utterance-level labels remains challenging due to scarce frame-level annotations. We propose P-SED, a weakly supervised SED framework based on prototype metric learning that constructs a metric space via learnable emotion prototypes and introduces orthogonal regularization to enhance inter-class separability. To mitigate sparse supervision and neutral background interference, we design a Class-aware Prototype Contrastive Loss (CPCL) combined with a Top-K ranking mechanism to mine salient emotion segments. Finally, Total Variation Denoising (TVD) is applied during inference to preserve steep emotion boundaries. Experiments on the ZED dataset show that P-SED significantly outperforms existing weakly supervised models, achieving an Emotion Diarization Error Rate (EDER) of 47.00%.
While Audio Language Models (ALMs) demonstrate strong semantic understanding, they struggle with complex affective interactions. Specifically, textual semantic dominance often overshadows acoustic nuances, and a lack of cognitive depth leads to generic, emotion-agnostic responses. We propose CogAudio-LLM, a novel cognitive affective reasoning framework. To mitigate semantic dominance, we build LIME-440K, a "lexically-identical, multi-emotion" dataset designed to facilitate acoustic-semantic decoupling. We introduce EIPS, a 4-step Chain-of-Thought (CoT) mechanism incorporating psychological reasoning. For inference efficiency, multi-stage training explicitly establishes EIPS via supervised fine-tuning, then distills this logic into an implicit generation process. Finally, we design DR-SAPO (Dual-Route Soft Adaptive Policy Optimization) to dynamically balance the logical rigor of the CoT with the empathetic quality of the direct response.
Current Large Audio Language Models (LALMs) can describe the sounds present in audio but cannot localize their occurrence-a critical gap for applications. Prior efforts to add temporal reasoning either rely solely on coarse multi-choice distinctions or depend on LLM-inferred timestamps whose accuracy is unverifiable. We address this with a time-aware audio instruction-tuning dataset, AudioGround-IT, that provides deterministic boundary supervision, yielding 49.9K instructions over 835 hours of audio across four temporal tasks. To leverage this supervision for temporal grounding in LALMs, we propose AudioGround, a lightweight extension of SALMONN that uses a sliding-window Q-Former to compress encoder features and receives timestamp conditioning and absolute time embeddings. Evaluations on multiple temporal grounding benchmarks show that AudioGround substantially outperforms prior LALMs, demonstrating that deterministic boundary supervision transfers effectively to real-world audio.
Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability central to human auditory perception. Understanding the causes of these failures remains challenging as existing benchmarks report performance gaps without probing underlying mechanisms. To address this, we introduce a benchmark with 1,657 questions across three foundational tasks designed specifically for mechanistic analysis. Examining model outputs across varying input settings (behavioral analysis) reveals that models often under-utilize audio when textual cues are available. We also provide the first causal mechanistic analysis of temporal reasoning failures in LALMs. Comparing attention upweighting against scaling, we find that redistributing attention across audio tokens is more effective than increasing audio attention. Targeting task-relevant tokens yields further gains. These findings suggest that modality imbalance alone cannot explain failures. Attention scaling at bottleneck layers improves accuracy from 55.9% to 59.1% without fine-tuning, demonstrating a promising direction for future work.
Speech recognition often fails on rare, domain-specific terms and context-related named entities. Existing contextualization techniques typically bias decoding with keywords or phrase lists, which does not scale well or exploit deeper knowledge. We propose a training method that teaches a speech-LLM to use broad descriptions (e.g. from videos) as weak semantic priors to perform contextual reasoning grounded in the audio. We build 400 hours of reasoning-augmented speech data by pairing erroneous hypotheses with video metadata and LLM-generated reasoning explanations that justify context-driven corrections. We finetune the speech-LLM to perform chain-of-thought reasoning: generate an initial transcript, then reason over the context, and finally return a corrected transcript. On held-out YouTube-derived test sets, our approach reduces errors, with specific improvements on rare words and named entities, and lays groundwork for deeper contextual reasoning in speech recognition.
Speech Large Language Models (SLLMs) underperform their text counterparts on complex reasoning. We reveal that this gap is not a uniform cognitive deficit. Evaluating two architecturally diverse SLLMs, we show speech-to-text (S2T) remains comparable to text-to-text (T2T) on spatial, syntactic, and factual tasks. Yet on logical tasks requiring entity tracking, S2T accuracy collapses to chance. We diagnose this as an entity binding failure: continuous speech features blur precise entity-property associations during implicit reasoning. To validate this diagnosis, we introduce Entity-Aware Chain-of-Thought (EA-CoT), a lightweight inference-time intervention forcing SLLMs to enumerate entities and bind them to claims before reasoning. EA-CoT bridges the gap, even when spoken names are misrecognized, yielding up to a 24.4 percentage-point accuracy gain. Ablations confirm the gains stem from explicit semantic binding, reframing the gap as an elicitation failure rather than a missing capability.
Auditory Large Language Models (ALLMs) show strong performance in audio understanding and reasoning, but their reliability is limited by hallucinations. Existing evaluation methods cast hallucination detection as binary classification, failing to capture the nuanced patterns in generative audio tasks, while mitigation methods often rely on costly fine-tuning. To address this, we propose a plug-and-play Noise-Aware In-Context Learning (NAICL) method. NAICL constructs a noise prior library, retrieves noise examples relevant to the input audio, and incorporates them as contextual priors to reduce speculative associations and generate more conservatively when acoustic evidence is insufficient. We also establish the Clotho-1K hallucination benchmark, define four types of auditory hallucinations, and introduce fine-grained evaluation metrics. Experiments show that evaluated ALLMs share similar hallucination behaviors, and NAICL reduces the hallucination rate from 26.53% to 16.98%.
Large Audio-Language Models (LALMs) achieve strong performance on multiple-choice Audio Question Answering (AQA) but often exhibit modality bias, over-relying on textual priors in questions and candidate options rather than grounded acoustic evidence. We present CoRE, a training-free, plug-and-play test-time option re-scoring method. CoRE constructs counterfactual audio via chunk permutation and random segment reversal to disrupt long-range temporal structure while largely preserving short-time acoustics. It estimates option-level evidence gain by contrasting scores from original and counterfactual audio, and applies an adaptive evidence-aware gate for final prediction. Under a unified option-scoring protocol, experiments on DCASE 2025 Task 5 and AIR-Bench SoundQA show consistent gains with Qwen2-Audio and Kimi-Audio.
Large Audio-Language Models (LALMs) excel in Audio QA but often suffer from hallucinations ungrounded in the audio. To our knowledge, we are the first to propose applying vector steering to the audio domain to mitigate this. Unlike text-based steering, our silence-anchored contrastive approach steers the model away from hallucinations by contrasting active audio against a silent baseline. Probing internal states reveals a strong correlation between specific layer representations and output correctness. Leveraging this, we introduce Layer-Weighted Vector Steering (LWVS), a training-free intervention that increases steering strength at influential layers. On the Audio Hallucination QA dataset, LWVS significantly outperforms baselines, boosting Recall on the Gemma model by 15.6% (53.4% to 69.0%). Crucially, MMAU benchmark tests confirm LWVS preserves and even enhances general audio understanding, achieving an 8% relative accuracy increase on the Qwen model (54.8% to 59.2%).
Streaming multi-speaker ASR is a challenging task that must balance accuracy, latency, and efficiency while handling overlapping speech and maintaining coherent longcontext modeling over extended conversations in an online fashion. We present a unified framework that categorizes streaming multi-speaker ASR into four architectural strategies based on how diarization and ASR are integrated. Using a shared pair of open-source streaming ASR and diarization models as a common foundation, we derive four multi-speaker ASR systems that differ in whether they employ multiple model instances, fine-tuning, or both. We evaluate these systems across multi-speaker accuracy, single-speaker accuracy degradation, memory footprint, and training complexity. Through this systematic architectural analysis, we clarify the design space for streaming multi-speaker ASR and provide practical guidance for selecting the most suitable approach under diverse deployment constraints.
Large Audio Language Models (LALMs) suffer from Neutral Bias, failing to capture subtle emotional nuances. Existing attempts to address this through text augmentation usually introduce acoustic hallucinations. To address these issues without costly human annotation, we propose a data-centric paradigm: Semantic Drift and Discriminative Re-ranking. Our method employs an LLM to perform multi-dimensional semantic drift, exploring diverse emotional hypotheses beyond neutral baselines. Subsequently, a discriminative judge model filters drifted hypotheses via hard-negative contrastive learning, grounding them in raw acoustic signals. Experiments confirm this hypothesize-and-verify mechanism transforms coarse, neutral-biased labels into precise, acoustically grounded descriptions, enriching fine-grained emotional detail and downstream expressiveness.
Data-efficient neural decoding is a central challenge for speech brain-computer interfaces. We present the first demonstration of transfer learning and cross-task decoding for MEG-based speech models spanning perception and production. We pretrain a Conformer-based model on 50 hours of single-subject listening data and fine-tune on just 5 minutes per subject across 18 participants. Transfer learning yields consistent improvements, with in-task accuracy gains of 1-4% and larger cross-task gains of up to 5-6%. Not only does pre-training improve performance within each task, but it also enables reliable cross-task decoding between perception and production. Critically, models trained on speech production decode passive listening above chance, confirming that learned representations reflect shared neural processes rather than task-specific motor activity.
Aligning non-invasive neural activity with speech stimuli is a foundational challenge in neural speech decoding. The difficulty lies in mapping noisy, high-dimensional electroencephalography (EEG) signals to the temporal dynamics of speech features. We propose SCANS, a supervised contrastive learning framework for neural-speech temporal alignment. We define this as a classification task where, given an EEG segment and multiple non-overlapping candidate segments from the same speech signal, the model must identify the single matching stimulus temporally aligned with the EEG. SCANS utilizes a dilated convolutional frontend and cross-modal attention to extract and fuse features across modalities. To bridge the modality gap, we employ a multi-task objective combining cross-entropy classification with a contrastive loss. Evaluated on the SParrKULee dataset, SCANS demonstrates significant improvement in neural-speech alignment accuracy.
Speech conveys not only linguistic information but also critical cues to speaker identity. Previous studies have identified bilateral superior temporal cortex (STG/STS) as central to human voice perception. However, speaker identity, comprising multiple-layer information, has largely been treated as a unitary construct. The present fMRI study investigated whether distinct speaker traits, such as gender, age, and accent, are processed by shared or dissociable neural mechanisms. Results revealed a common voice-processing core centered on bilateral STG/STS, which was engaged across all conditions. However, neural responses within this core were not uniform, with accent processing eliciting stronger activation than gender and age. Beyond this shared region, each trait recruited partially distinct cortical regions. These findings revealed a "core-plus-extension" model of speaker identification, in which shared auditory mechanisms are complemented by trait-specific cortical regions.
Speaker normalization is essential for speech perception, particularly in tonal languages like Cantonese. Listeners can normalize talker variability under cognitive load, yet the neural mechanisms remain unclear. We recorded EEG from Cantonese speakers during a tone normalization task with and without concurrent visual tasks. Although behavioral performance was unaffected by visual distraction, neural data revealed distinct processing strategies. Without load, alpha-band desynchronization emerged, but load conditions triggered sustained parieto-occipital alpha activation, reflecting active suppression of visual distractors. High load uniquely induced prefrontal delta-band inhibition, indicating an attentional shift from internal visual processing to external auditory stimuli. These results suggest that normalization under load is not automatic. Instead, the brain actively reallocates resources to shield speech processing from distraction, supporting an active control hypothesis.
The Post-Movement Beta (13-30Hz) Rebound (PMBR) is a motor cortex specific response that follows the termination of a movement, reflecting various aspects of performance. While widely researched in limb-movement tasks, the PMBR has not been extensively studied in speech. This study used magnetoencephalography (MEG) to characterize the PMBR response in five healthy adults as they spoke five phrases varying in articulatory complexity. Source localization was conducted using beamforming, and the PMBR was evaluated with respect to jaw movement offset. Groupwise results revealed that the most complex phrase elicited a stronger PMBR than the least complex phrase. The PMBR was found bilaterally, but with a stronger response in the language-dominant hemisphere. These findings provide preliminary evidence that the PMBR could be used as an index of articulatory complexity during speech production. The PMBR may even serve as a useful neurophysiological feature for clinical speech applications.
Speech recognition in competing masker speech can be influenced by the similarity between the target and masker languages, and the listener's familiarity with them. Generally, recognition is more difficult when the target and masker are identical or similar languages. While greater familiarity with the target language is beneficial, familiarity with the masker can be detrimental. In this study, Australian English monolinguals and Arabic–English bilinguals completed an English word monitoring task presented with different masker languages. The English masker led to the poorest accuracy and response times, but there was no evidence that the degree of typological similarity between the target and masker also affected word monitoring. The Arabic–English bilinguals, while slower, were not less accurate than the monolinguals, nor were they hindered by the Arabic masker. Cognitive load, which was manipulated by a digit preload secondary task, did not seem to affect word monitoring outcomes.
We examined whether distinct forms of bilingual language use are differentially associated with executive function (EF) in young adults. Shifting, working memory, and inhibition were modelled as correlated latent variables. Motivations for code-switching were derived via exploratory factor analysis. Overall switching frequency was unrelated to EF. Lexically motivated switching was positively associated with updating and inhibition, whereas frequent language brokering was negatively associated with these domains. Results suggest that associations between bilingual speech and EF depend on communicative function rather than switching frequency alone.