| Total: 1375
Recent breakthroughs in multi-talker ASR (MT-ASR) and speaker diarization (SD) rely on synthetic data to mitigate the scarcity of large-scale conversational recordings, yet the impact of specific simulation choices remains poorly understood. To mind the gap between simulated mixtures and real-world interactions, we present a study of synthetic data generation for leading MT-ASR (DiCoW) and SD (Sortformer) systems. By introducing FastMSS, a highly efficient open-source simulator, we analyze turn-taking dynamics, source domain, acoustic augmentation, and data mixing strategies. Our findings reveal that optimal simulation recipes are highly task-dependent: increasing speech overlap benefits ASR but degrades diarization. Furthermore, broad source diversity consistently outperforms exact domain matching. Ultimately, synthetic-only training approaches real-data baselines, and combining simulated data with real recordings yields substantial gains over real-only training across both tasks.
We propose diarization-conditioned spoken language models (SLMs), a strategy for extending SLMs to far-field multi-talker audio. Rather than adapting the decoder via Serialized Output Training, which risks catastrophic forgetting, we condition the acoustic encoder on diarization masks to extract target-speaker representations, keeping the decoder frozen. We instantiate this as Dixtral, integrating a Diarization Conditioned Whisper (DiCoW) encoder into the Voxtral SLM. On AMI, NOTSOFAR-1, LibriSpeechMix, and Mixer6, Dixtral outperforms Gemini 3.0 Flash, VibeVoice, and Voxtral Mini Transcribe V2 on speaker-attributed transcription by 29.0%, 19.8%, and 16.0% absolute cpWER respectively. On a novel long-form multi-speaker QA benchmark, zero-shot Dixtral matches Gemini on far-field content understanding, and when fine-tuned surpasses both Gemini and Voxtral operating on close-talk across all tasks.
End-to-end multi-talker automatic speech recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech. A critical bottleneck is that speaker-specific acoustic characteristics, which are essential for distinguishing overlapping speech, are often diluted in deep network layers. To address this, we propose the Global-Local Aware Dynamic Mixture-of-Experts (GLAD) architecture. GLAD introduces a novel routing mechanism that dynamically fuses speaker-aware global context with fine-grained local acoustic details to adaptively guide expert selection. Experiments on the LibriSpeechMix and CH109 datasets demonstrate that GLAD significantly outperforms existing Serialized Output Training (SOT)-based MTASR approaches, exhibiting exceptional robustness in challenging, high-overlap scenarios. To the best of our knowledge, this is the first work to apply a global-local fusion MoE strategy to MTASR.
Multi-talker automatic speech recognition (ASR) has attracted increasing attention for overlapping speech scenarios. Hypothesis Clustering and Merging (HCM) achieves strong performance by clustering hypotheses in transcript space, but it does not explicitly consider speaker identity during clustering. As a result, HCM may fail when multiple speakers utter identical content and may not fully utilize enrollment information in target-speaker settings. In this paper, we incorporate continuous speaker embeddings into the HCM framework by redefining the clustering distance in joint transcript-speaker space. The proposed method improves robustness in target-speaker-free scenarios under identical-content conditions and enables more effective use of enrollment information in target-speaker multi-talker ASR. Experimental results show consistent improvements over conventional HCM, achieving up to 46% relative WER reduction under identical-content conditions while maintaining competitive performance on LibriMix benchmarks. In addition, the proposed speaker-aware formulation improves target-speaker multi-talker ASR by enabling embedding-based speaker selection, achieving 23% WER reductions compared to discrete speaker-ID prompting.
Streaming multi-speaker ASR is a challenging task that must balance accuracy, latency, and efficiency while handling overlapping speech and maintaining coherent long-context modeling over extended conversations in an online fashion. We present a unified framework that categorizes streaming multi-speaker ASR into four architectural strategies based on how diarization and ASR are integrated. Using a shared pair of open-source streaming ASR and diarization models as a common foundation, we derive four multi-speaker ASR systems that differ in whether they employ multiple model instances, fine-tuning, or both. We evaluate these systems across multi-speaker accuracy, single-speaker accuracy degradation, memory footprint, and training complexity. Through this systematic architectural analysis, we clarify the design space for streaming multi-speaker ASR and provide practical guidance for selecting the most suitable approach under diverse deployment constraints.
Conversational automatic speech recognition remains challeng-ing due to overlapping speech, far-field noise, and varying speaker counts. While recent LLM-based systems perform well on single-speaker benchmarks, their robustness in multi-speaker settings is unclear. We systematically compare LLM-based and modular pipeline approaches along four axes: overlap robustness, semantic fidelity speaker count, and single- versus multi-channel input. To capture meaning-altering errors that conventional metrics miss, we introduce tcpSemER, which extends tcpWER by replacing Levenshtein distance with embedding-based semantic similarity. We further decompose tcpWER into overlapping and non-overlapping components for finer-grained analysis. Experiments across three datasets show that LLM-based systems are competitive in two-speaker settings but degrade as speaker count and overlap increase, whereas modular pipelines remain more robust.
In sequence-to-sequence Transformer-based automatic speech recognition (ASR), autoregressive (AR) models achieve strong accuracy but suffer from slow decoding, while non-autoregressive (NAR) models enable parallel decoding at the cost of degraded performance. We propose a principled NAR ASR framework based on Masked Diffusion Models to significantly reduce this gap. A pre-trained speech encoder is coupled with a Transformer diffusion decoder conditioned on acoustic features and partially masked transcripts for parallel token prediction. To mitigate the training-inference mismatch, we introduce Iterative Self-Correction Training that exposes the model to its own intermediate predictions. We also propose a Position-Biased Entropy-Bounded Confidence-based sampler to further boost results. Experiments across multiple benchmarks demonstrate consistent gains over prior NAR models and competitive performance with strong AR baselines, while retaining parallel decoding efficiency.
Whisper, a widely adopted ASR model, is known to suffer from hallucinations -- coherent transcriptions generated for non-speech audio entirely disconnected from the input. We investigate whether hallucinations can be detected and mitigated through Whisper's internal representations. We extract audio encoder activations and evaluate two representation spaces: raw Whisper activations and Sparse AutoEncoder (SAE) latents. We show that both spaces encode linearly separable hallucination-related information, with discriminative power concentrated in a sparse feature subset and increasing toward deeper encoder layers. We propose two steering strategies: activation-space steering and SAE latent-space steering. SAE-based steering reduces hallucination rate from 72.63% to 14.11% for Whisper small and from 86.88% to 27.33% for Whisper large-v3 on the full non-speech test set, with small WER degradation on speech data, approaching the performance of fine-tuning-based methods.
Modern ASR systems are typically trained on large-scale pseudo-labeled, in-the-wild data spanning multiple domains. While such heterogeneous data benefit generalist models designed for broad deployment, they pose challenges for specialist models targeting specific domains: specialist models lack the capacity to learn from all available data. In this work, we study targeted data selection to address these challenges, selecting relevant subsets from 100k hours of in-the-wild training data to optimize performance on target domains. We represent speech samples using embeddings that capture complementary characteristics—speaker attributes, phonetic content, and semantic meaning—and study how relevance and diversity along these axes when performing data selection affect ASR performance. Our experiments with CTC-based models show that training on a strategically selected 5% subset can exceed the performance of models trained on the full data by up to 36.8% relative WER reduction.
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER from style mismatch), and unreliable word-level timing. We show models already encode both styles, the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot from English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We introduce supervised cross-attention finetuning improving word-level timestamps on disfluent speech beyond forced alignment baselines. Finally we introduce a new task: verbatimize, enabling scalable creation/enrichment of speech corpora with high quality canonical verbatim transcriptions.
Streaming Inverse Text Normalization (ITN) is vital for converting spoken-form outputs from streaming Automatic Speech Recognition into formatted written text. Existing streaming ITN methods rely on hybrid systems combining neural tagging with expert-crafted finite-state transducer rules, limiting scalability across domains and languages. While end-to-end models offer superior scalability by learning directly from data, standard encoder-decoder architectures are inherently non-streaming due to global attention. In this paper, we propose an efficient streaming end-to-end ITN system, adapting a pretrained text-to-text model to leverage its robust linguistic knowledge. To enable streaming, we introduce architectural adaptations, a specialized training strategy, a Read-Tag-Write decoding policy, and inference optimizations. Experiments on Vietnamese datasets show accuracy comparable to non-streaming baselines, outperforming hybrid methods while satisfying real-time latency requirements.
Test-Time Adaptation (TTA) via entropy minimization (EM) has proven effective for classification tasks, yet its application to generative autoregressive models remains theoretically fragmented. Existing approaches typically rely on distinct heuristics, such as teacher forcing with pseudo labels or policy-gradient-based reinforcement learning, without a unified mathematical foundation. In this work, we resolve this discrepancy by deriving a rigorous formulation of EM tailored to autoregressive models. We show that the exact objective naturally decomposes into a token-level policy gradient loss and a token-level entropy loss, and we reinterpret prior methods as partial realizations of this unified formulation. Using Whisper ASR as a testbed, we demonstrate that our approach consistently improves performance across more than 20 diverse domains, including acoustic noise, accents, and multilingual settings.
Speech emotion recognition (SER) faces challenges due to label ambiguity and high model uncertainty during early training. Standard training with hard one-hot targets ignores these issues, forcing models to commit to a single class before learning meaningful representations. We propose Progressive Weak Supervision (PWS), a strategy that relaxes supervision in the early stages and gradually tightens it as the model matures. During the initial epochs, a sample is considered correctly predicted if the true label is among the top-k predictions, with a soft probability distribution assigned across them. As training progresses, k decays linearly to 1, thereby recovering standard supervision. We evaluate PWS using a WavLM-Base encoder on IEMOCAP (English, 4-class) and ViSEC (Vietnamese, 4-class). With k=3, PWS achieves 78.08% unweighted accuracy on IEMOCAP and 85.70% on ViSEC, surpassing several strong baseline methods. Code is available at https://github.com/skyemo47/PWS.
Speech Emotion Recognition (SER) benefits from large Self-Supervised Learning (SSL) models (e.g., HuBERT, WavLM), but their immense computational overhead hinders real-time edge deployment. While mel-spectrograms offer a lightweight alternative, traditional CNNs struggle to match SSL representation quality. To bridge the gap between the two approaches, we propose MSMC (Multi-Scale Masked Convolution), a highly efficient architecture for spectrogram-based SER. A masked convolution encoder (MCE) extracts sparse spectral features without information leakage, while a mean teacher framework enforces multi-scale consistency, to effectively distil global semantic context and micro-prosodic details into a compact model. Evaluated on the IEMOCAP dataset, MSMC achieves state-of-the-art accuracy among lightweight models (76.0% WA), matching the performance of heavy SSL baselines. With a significantly lower parameter count and computational complexity, MSMC is suitable for real-time SER applications.
Self-supervised learning (SSL) yields powerful, context-rich representations for speech emotion recognition (SER), yet the aggregation of these representations into holistic descriptors remains a bottleneck. Conventional first-order aggregation implicitly assumes feature independence, violating the latent Riemannian geometry and discarding higher-order relationships essential to the backbone's representational power. To address this problem, this paper proposes a novel Second-Order Correlation (SOC) layer. Instead of treating features in isolation, our SOC models correlations among features as covariance descriptors to capture synergistic co-occurrence patterns, which act as discriminative signatures for robust emotion recognition. By mapping these descriptors from the Riemannian manifold to a Euclidean tangent space through Log-Euclidean mapping (LEM), our method preserves geometric integrity while enabling direct linear discriminative learning. Extensive experiments on ESD and RAVDESS datasets demonstrate that SOC recovers discriminative information lost in first-order pooling, effectively aggregating high-dimensional SSL features.
Speech Emotion Recognition (SER) models and Audio LLMs fail when vocal tone contradicts textual semantics. Extensive training on congruent data engenders a rigid semantic prior, causing "Semantic Dominance", a severe text bias during cross-modal conflicts. We introduce the ASPIRE benchmark of adversarial audio-text pairs across four conflict types, alongside two metrics: Semantic Overconfidence Penalty (SOP) and Latent Decoupling Degree (LDD). Unlike fusion methods that merely re-weight collapsed features, our Acoustic Conflict Resolution Network (ACR-Net) uses Cross-Modal Attention and Contrastive Decoupling Loss to disentangle contradictory representations into orthogonal latent spaces. Experiments show existing models fail under polarity conflicts due to extreme semantic overconfidence. Conversely, ACR-Net preserves acoustic fidelity, maintaining high LDD to drive SOP near zero, achieving superior Acoustic Accuracy (ACC) in both adversarial and standard scenarios.
Post-training quantization (PTQ) methods such as GPTQ and AWQ compress large language models effectively, but applying them to automatic speech recognition (ASR) at ultra-low bit-widths (2-3 bits) causes severe transcription degradation. Unlike discrete text tokens, speech activations are highly correlated in steady-state and zero-padded regions yet change rapidly at phonetic boundaries; when all frames contribute equally to the Hessian, static frames dominate calibration and obscure informative transitions. We propose DiffAQ, which computes frame-to-frame activation differences to measure the rate of acoustic change and assigns Hessian importance proportionally, concentrating quantization precision on phonetically critical frames. As a training-free modification to GPTQ, DiffAQ consistently reduces WER across various Whisper sizes and standard benchmarks, with the largest gains at 2-bit precision where baseline methods frequently produce degenerate outputs.
Memristors provide a new chance for resource-efficient computation of neural models for natural language processing by enabling analog execution of vector-matrix-multiplication. Yet, computations on these devices are currently subject to larger distortion, both in weight programming and execution. In this work, we identify large output values of transformed positional encodings to cause major degradation within analog-to-digital conversion (ADC) as part of memristor-based computation. By adjusting the proportion of weight and precision bits of the ADC of specific memristor layers, we reduce the degradation of the execution by ~50% relative, while keeping the estimated energy consumption stable. Additionally, we investigate scenarios where the ADC cannot be modified. In that case the degradation can be reduced by ~30% relative after removing encoding-related linear transformations.
Whisper and similar encoder-decoder ASR models are increasingly deployed on mobile and edge devices, yet it remains unclear how quantization format choices affect their accuracy in practice. We systematically evaluate post-training quantization (PTQ) for Whisper (tiny.en,base.en) across 80+ configurations covering INT8/4/3 and FP8/FP4/NVFP4/MXFP4. Our main finding is that activation bit-width matters far more than weight format: dropping activations from 16-bit to 8-bit costs 1-3% absolute WER, while INT16 and FP16 activations are indistinguishable in accuracy. Among 4-bit formats, NVFP4 W4A16 comes within 0.07% of full-precision at 6.4× compression, whereas MXFP4 fails under standard PTQ and needs additional fixes. FP activation paths are also the better hardware choice, since FP multipliers are more area-efficient than INT at the same bit-width. We close with a Pareto analysis and six practical guidelines for format selection across memory budgets from 20 to 80 MB.
While Conformer-Transducers offer state-of-the-art ASR performance, their excessive memory footprints create a bottleneck for on-device deployment. Conventional integer quantization is fundamentally constrained by a discrete grid, imposing a structural lower bound of 1 bit per parameter. To address this challenge, we propose LittleASR, a mixed-precision framework based on variable-rank binary decomposition. Guided by a gradient-aware sensitivity metric, LittleASR identifies and compresses non-critical layers to the sub-1-bit regime, pushing the limits of compression beyond the boundaries of standard quantization while maintaining a favorable trade-off between extreme memory reduction and recognition accuracy.
Neural network pruning is typically framed as a post-training compression technique. We show that for encoder-decoder ASR Transformers, one-shot magnitude pruning can act as a strong implicit regularizer: removing redundant parameters improves generalization without fine-tuning. Using Whisper-small, we introduce a sensitivity diagnostic combining gradient and Fisher criteria to identify pruning-fragile vs. pruning-resilient components. This reveals an encoder-decoder asymmetry: decoder FFNs are pruning-fragile, whereas decoder self-attention and late encoder layers contain removable redundancy. Without fine-tuning, pruning 50% of decoder self-attention improves WER by 2.38% absolute on LibriSpeech test-other; pruning the last four encoder layers at 50% yields 1.72% improvement. Gains persist on Common Voice and TED-LIUM, demonstrating cross-corpus generalization. At 40.8% sparsity, sensitivity-aware compression preserves near-baseline accuracy where global magnitude pruning collapses.
Transformer-based models have revolutionized Automatic Speech Recognition. However, the quadratic complexity of self-attention and feature redundancy necessitate deep stacking, limiting long-sequence efficiency. To address this, we propose Key-Value Pooling Attention (KV-Pooling), which leverages high temporal redundancy in speech by average pooling to compress Key/Value tensors, smoothing feature distribution spikes while retaining Query resolution to maintain modeling precision. This reduces reliance on network depth. Building on this, we introduce KV-Pooling-Zipformer, integrating this mechanism into Zipformer by pre-setting differential pooling strides based on layer-wise feature abstraction. Compared to the native Zipformer, RNN-T experiments show absolute reductions of 0.2% CER on AISHELL-1 and 0.3% WER on LibriSpeech. Notably, it improves inference RTF by 10% on an AMD EPYC 7763, achieving synergistic gains in accuracy and efficiency.
Speech Emotion Recognition (SER) plays a key role in advancing human-computer interaction. Attention mechanisms have become the dominant approach for modeling emotional speech due to their ability to capture long-range dependencies and emphasize salient information. However, standard self-attention suffers from quadratic computational and memory complexity, limiting its scalability. In this work, we present a systematic benchmark of optimized attention mechanisms for SER, including RetNet, LightNet, GSA, FoX, and KDA. Experiments on both MSP-Podcast benchmark versions show that while standard self-attention achieves the strongest recognition performance across test sets, efficient attention variants dramatically improve scalability, reducing inference latency and memory usage by up to an order of magnitude. These results highlight a critical trade-off between accuracy and efficiency, providing practical insights for designing scalable SER systems.
We present a reproducible multi-label benchmark for speech emotion recognition (SER) built from the MSP-Podcast V2.0 corpus. Our pipeline converts crowdsourced emotion votes into multi-hot labels with controlled ambiguity and sparsity via reliability filtering, deterministic vote-to-label construction, and lightweight time-shift augmentation. On this benchmark, we provide a compact Mamba-based fusion baseline over MFCC and log-mel features, following prior SER fusion practice, while using linear-time state-space backbones for efficient modeling of long audio. Using a validation-only global threshold sweep for checkpoint selection and a fixed test-time threshold for evaluation, our fusion model attains competitive micro-F1 on two MSP-Podcast test partitions (≈0.50 on both Test1 and Test2) while remaining substantially lighter than heavier baselines and recent methods. We release dataset splits, construction code, and training recipes to support future work on scalable multi-label SER.
Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce annotated data. To tackle these challenges, we propose a multi-loss learning (MLL) framework integrating an energy-adaptive mixup (EAM) method and a frame-level attention module (FLAM). The EAM method leverages SNR-based augmentation to generate diverse speech samples capturing subtle emotional variations. FLAM enhances frame-level feature extraction for multi-frame emotional cues. Our MLL strategy combines Kullback-Leibler divergence, focal, center, and supervised contrastive loss to optimize learning, address class imbalance, and improve feature separability. We evaluate our method on four widely used SER datasets: IEMOCAP, MSP-IMPROV, RAVDESS, and SAVEE. The results demonstrate our method achieves state-of-the-art performance, suggesting its effectiveness and robustness.