| Total: 373
Automatic speech recognition systems have been shown to under-perform when it comes to transcribing words rarely seen in the training data, namely specialized terminology. Open-vocabulary keyword spotting, combined with contextual biasing, has been shown to mitigate this issue. However, existing systems can only handle glossaries of a few hundred terms without becoming an infeasible bottleneck. We propose a system that stores features with a memory footprint up to 128 times smaller than a comparable baseline and allows users to process massive databases while remaining open-vocabulary. Without fine-tuning the speech recognition model, our system achieves a comparable entity recall as uncompressed solutions, even in languages not seen during training.
With the increasing prevalence of voice-driven interaction, keyword spotting (KWS) has become an essential component of hands-free control. This has led to a growing demand for user-defined KWS, allowing users to customize target keywords via text. While various models have emerged to support this, their high computational costs and energy consumption make real-world deployment challenging. In response, we introduce SPARK, a SPike-driven Audio-text matching framework for eneRgy-efficient user-defined Keyword spotting. SPARK employs a spike-driven attention mechanism to enable end-to-end processing within the spiking domain, replacing heavy floating-point operations with low-cost accumulate operations. Our experimental results on the LibriPhrase test dataset demonstrate that SPARK achieves competitive performance with a 2.1 times reduction in parameter count and a 21.7 times reduction in energy consumption compared to its artificial neural network counterpart.
Keyword spotting (KWS) models on embedded devices often need to add new keywords after deployment, but updates are difficult when original training data are unavailable and regressions on existing triggers are unacceptable. At a fixed operating point, our method reduces average new-keyword false reject rate (FRR) from 6.46% to 4.37% versus a parameter-matched separate-model baseline and outperforms parameter-efficient tuning baselines (adapters, LoRA), while using fewer multiply-accumulate operations (MACs) under the same added-parameter budget (≤10k): 16.34M vs 18.45M/20.52M. We achieve this via parameter-capped modular expansion: the base network, including batch-normalization statistics and the core classifier, is frozen, and only a lightweight expansion branch with a separate new-keyword head is trained, preserving core logits, shipped outputs, and thresholds for existing keywords.
In recent years, attention-based multi-modal open-vocabulary keyword spotting (KWS) has attracted attention, yet its streaming deployment on device with such a design remains underexplored. We identify a data flow mismatch in traditional frameworks: speech as Key/Value requires global context but, in streaming, provides only local frames; enrollment representation as Query supplies only local information, despite its inherent global semantics. To resolve this mismatch, we propose a role-swapping streaming open-vocabulary KWS approach: streaming speech becomes Query (carrying frame-level local info) and enrollment representation as Key/Value (encoding global semantics), enabling streaming processing. Without loss of generality, we instantiate this framework with a text-registered model (~0.8M parameters) and evaluate it on LibriPhrase, achieving an EER/AUC of 6.82%/97.95% on easy negative subset and 28.21%/79.19% on hard negative subset.
For open-vocabulary keyword spotting, phoneme-level alignment has been introduced to improve the performance on acoustically confusable words. However, most existing studies utilize non-streaming methods, which are unsuitable for streaming scenarios. Recent methods achieve streaming phoneme alignment via connectionist temporal classification (CTC) algorithms, but they are limited to text-only enrollment. To address these issues, we propose a novel streaming multimodal open-vocabulary keyword spotting method based on phoneme-level alignment. This method achieves fine-grained modeling with training-inference consistency through the W-CTC forced alignment algorithm and multimodal phoneme-level contrastive learning. Additionally, we propose a data augmentation method based on CTC beam-search to dynamically mine hard negative samples. Experiments on the LibriPhrase dataset demonstrate that the proposed method achieves the best results.
This paper investigates audio-to-audio retrieval using self-supervised learning (SSL) models to generate audio representations without labeled data. To enhance retrieval accuracy, we explore the use of SSL embeddings with sequence matching techniques, including Dynamic Time Warping (DTW), and clustering methods, such as K-Means combined with TF-IDF and BM25. We evaluate our framework on two distinct tasks: music retrieval via query-by-humming and spoken content retrieval via query-by-example. Experimental evidence shows that clustering-based methods, which reduce SSL embeddings into discrete hidden units, are particularly effective for speech retrieval. Conversely, DTW applied directly on SSL embeddings, which preserves full sequence information, excels in music retrieval. Extensive experiments demonstrate that combining SSL representations with appropriate sequence matching improves retrieval accuracy across different audio domains.
When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may be heard but unseen, seen but unheard, or both. Fusing scores from an absent modality injects noise, degrading precision below the best unimodal system. We propose a query-adaptive framework that detects active modalities via cross-modal score consistency: when both modalities are active, files retrieved by one also score highly on the other; this agreement breaks down when a modality is absent. Classifiers driven by these cross-modal features achieve 89% detection accuracy. On the BBC Rewind corpus (with over 12,000 broadcast videos) the adaptive system attains 94.2% P@1, outperforming speaker-only (82.9%), face-only (93.4%), and fixed fusion (90.0%), recovering 64% of the gap to an oracle with ground-truth modality labels (96.6%).
Detecting depression from speech in elderly patients with mild cognitive impairment (MCI) is complicated by overlapping acoustic effects of cognitive decline. Without disentangling these, classifiers risk learning cognitive rather than depression-specific patterns. We present a Korean elderly speech corpus of 89 MCI speakers collected over three years with concurrent depression (SGDS) and cognitive (MMSE) assessments. Both linear mixed-effects models and partial correlations identify the same pattern: of seven feature groups, only formant features are associated with depression after controlling for cognitive function, while widely used F0 features carry no signal. The F1+F2 subset achieves UAR 0.760 with just 12 features, outperforming both the full eGeMAPS baseline and self-supervised representations. These findings show that concurrent clinical assessments can identify interpretable, depression-specific acoustic markers in populations where cognitive and affective symptoms co-occur.
Speech-based cognitive impairment detection offers a noninvasive, accessible alternative to costly biomarker assays, yet transformer-based models remain clinically uninterpretable. We propose a multi-stage explainability framework that translates black-box transformer predictions into clinically grounded narratives by integrating SHapley Additive exPlanations (SHAP)-based token attribution, theory-informed linguistic features, and a four-stage LLM reasoning pipeline using LLaMA-3.1-70B-Instruct. Built on the SpeechCARE-Adaptive Gating Network multimodal screening model (F1 = 72.11% on the NIA PREPARE benchmark), the framework maps model outputs to four cognitive-linguistic dimensions, including lexical richness, syntactic complexity, and semantic coherence. Physician evaluation on 70 stratified English samples demonstrated strong alignment with patient-level cognitive profiles, and a System Usability Scale score of 82/100 indicated high potential for clinical workflow integration.
Amyotrophic Lateral Sclerosis (ALS) is a neurodegenerative disease, often affecting speech due to bulbar dysfunction. In this study, we predict speech impairment in people with ALS (pwALS) using two clinical speech-related scores. We evaluate cross-sectional (across speakers) and personalised (within-speaker) modelling paradigms and analyse the utility of common speech tasks to contribute to the standardisation of speech data collection for pwALS. Experiments on a German-speaking cohort of 66 pwALS show that repetition tasks (/da/-/da/, /da/-/ba/) achieved the best cross-sectional performance (Concordance Correlation Coefficient (CCC) = 0.62) for predicting the Quality of Life in the Dysarthric Speaker questionnaire, while the within-speaker setting reached a CCC of 0.86. This study represents an initial step towards speech impairment prediction in German-speaking pwALS and highlights the potential of automated speech analysis as a supportive tool for speech impairment assessment.
The Cookie Theft picture description task is widely used to assess cognitive–linguistic abilities. Analyzing how speakers progress through the 23 content information units (CIUs) in the picture provides insight into the informational relevance and efficiency of their descriptions. Although prior CIU-based studies have shown effectiveness in distinguishing cognitively impaired speakers, they largely rely on manual annotation and focus on spatial CIU distributions, leaving temporal narrative dynamics underexplored. We propose an automated framework for identifying CIUs from picture descriptions and modeling transitions between CIUs as a temporal graph. Graph-based features are extracted to characterize narrative organization and temporal dynamics, and a visualization is designed to illustrate how speakers traverse picture content over time. The framework extends CIU-based analysis beyond spatio-semantic representations and offers a scalable approach for assessing cognitive abilities.
Alzheimer's disease and related dementia (ADRD) remain largely undiagnosed, as early cognitive symptoms are rarely captured by structured clinical data. Natural speech offers a sensitive, non-invasive window into cognitive decline that structured clinical data cannot provide. This study validates conversational speech from phone calls and participant verbal communication as a biomarker for early cognitive impairment detection. Using an attention-based fusion model across 175 participants, structured clinical data alone achieved AUC = 0.74; adding phone calls improved performance (AUC = 0.76), and participant verbal communication yielded the largest gains (AUC = 0.90). The best configuration achieved AUC = 0.92 and F1 = 83.08, establishing naturalistic speech as a scalable, non-invasive biomarker in clinical settings.
Early detection of functional decline in amyotrophic lateral sclerosis (ALS) is critical for timely intervention and efficient clinical trial design. We evaluated speech-derived biomarkers to detect ALS-related functional decline events earlier than traditional clinical measures. Using longitudinal patient data, we applied Kaplan-Meier analysis to estimate time-to-event and derived hazard rates for each measure. Our results demonstrate that speech-based biomarkers identify functional decline faster than conventional ALS Functional Rating Scale - Revised (ALSFRS-R) subscores. Hazard rate modelling enabled the estimation of sample sizes and trial durations required to detect treatment effects. These findings suggest that speech-based measures can accelerate ALS clinical trials by providing sensitive and objective endpoints, supporting more efficient study design and patient stratification.
The ADReSS and ADReSSo challenge datasets have become the de-facto standards for research on dementia detection through speech, with over half of recent ICASSP and Interspeech studies on the topic relying on them. Despite their widespread adoption, both datasets exhibit properties that may undermine the validity of reported results. In this paper, the acoustic variability in the test sets of ADReSS and ADReSSo is systematically examined. Results show that (i) near state-of-the-art classification performance on the test sets can be achieved using only two low-level acoustic features, (ii) classifiers trained on randomly permuted dementia labels can nonetheless reach competitive test performance, and (iii) feature discriminability does not generalize across Monte Carlo resampling, with no stable features emerging from low-level acoustic feature sets. These findings indicate that strong results on these datasets can stem from spurious correlations rather than pathology-relevant cues. The analysis calls for more cautious interpretation of benchmark performance, as well as renewed attention to small dataset evaluation strategies in dementia detection.
This study examines the relationship between speech representations and the hierarchical structure of cognitive assessment in mild cognitive impairment. Utilizing 5,754 German neuropsychological assessment recordings, we evaluate six cognitive tasks across three score levels: task, domain, and global levels. We compare hand-crafted acoustic features with self-supervised learning (SSL) embeddings. Results show that although SSL representations generally outperform hand-crafted features at lower levels, this trend reverses for MCI classification. Furthermore, task-specific constraints influence performance: tasks with greater response freedom exhibit performance dilution as hierarchical levels increase, suggesting "specialist" representations, whereas the performance of highly structured tasks increases toward higher levels, suggesting "generalist" representations. These findings show links between task constraints and assessment hierarchy in automated clinical speech analysis.
Spoken dialogue models typically start from text LLM backbones, yet reasoning often degrades when conditioning on speech instead of text. We attribute part of this modality gap to a temporal-granularity mismatch: speech tokens are temporally redundant and far longer than text under matched semantics, diluting per-token semantic density and weakening text-native reasoning dynamics. We study speech token design as a representation selection problem and sweep frame rates under a frozen LLM backbone with a fixed information rate. To make low frame rates feasible, we introduce factorized FSQ and a lightweight non-autoregressive audio LM head, scaling capacity to nearly 300 bits/frame without sacrificing efficient prediction. With the bottleneck removed, we sweep frame rates (50→2.08 Hz) and alignment depth, and observe a consistent best regime for speech QA at 4.17 Hz with intermediate-layer representation alignment.
Self-supervised learning (SSL) models, such as Wav2Vec2, HuBERT, and WavLM, have become foundational across a wide range of speech and audio tasks. Despite their success, understanding their internal layer-wise dynamics remains an ongoing challenge. To address this, we propose a two-part model-centric framework called INSIDESSL. First, we establish a task-agnostic analysis from three intrinsic per-layer perspectives: compression (entropy), geometry (curvature), and robustness to perturbations. We show that varying training objectives induce distinct regimes of acoustic compression and manifold unfolding. Second, we introduce the cross-layer Generative Compatibility Matrix (GCM) to evaluate functional transferability, exposing stable phonetic cores, identity volatility, and deep-layer semantic pruning. In addition to these evaluations, linear probing connects the model-centric perspective to downstream tasks, demonstrating how layer topology dictates phoneme, pitch, and speaker encoding.
The Montreal Forced Aligner (MFA) was released in 2016 and has since become the most widely used tool for forced alignment in research and industry. In the decade since, MFA has undergone substantial development, including expanded coverage across more languages and dialects using larger open-source datasets, harmonized IPA dictionaries, model adaptation, cross-language phone remapping, and support utilities. This paper documents MFA 3.0's developments since version 1.0 and evaluates MFA's performance across English, Japanese, and Korean, benchmarked against classic and neural forced aligners. MFA 3.0 achieves state-of-the-art or near state-of-the-art performance across all four benchmark datasets with mean boundary errors below 15 ms. Adaptation and cross-language remapping are effective for languages outside MFA's training distribution, and pronunciation probability modeling and phonological rules provide gains in specific conditions.
Self-supervised speech models (S3Ms) achieve strong downstream performance, yet their learned representations remain poorly understood under natural and adversarial perturbations. Prior studies rely on representation similarity or global dimensionality, offering limited visibility into local geometric changes. We ask: how do perturbations deform local geometry, and do these shifts track downstream automatic speech recognition (ASR) degradation? To address this, we present GRIDS, a framework using Local Intrinsic Dimensionality (LID) across layer-wise representations in WavLM and wav2vec 2.0. We find that LID increases for all low signal-to-noise ratio (SNR) perturbations and diverges at high SNR: benign noise converges toward the clean profile, while adversarial inputs retain early-layer LID elevation. We show LID elevation co-occurs with increased WER, and that layer-wise LID features enable anomaly detection (AUROC 0.78-1.00), opening the door to transcript-free monitoring in S3Ms.
Large audio language models (ALMs) extend LLMs with auditory understanding. A common approach freezes the LLM and trains only an adapter on self-generated targets. However, this fails for reasoning LLMs (RLMs) whose built-in chain-of-thought traces expose the textual surrogate input, yielding unnatural responses. We propose self-rephrasing, converting self-generated responses into audio-understanding variants compatible with RLMs while preserving distributional alignment. We further fuse and compress multiple audio encoders for stronger representations. For training, we construct a 6M-instance multi-task corpus (2.5M unique prompts) spanning 19K hours of speech, music, and sound. Our 4B-parameter ALM outperforms similarly sized models and surpasses most larger ALMs on related audio-reasoning benchmarks, while preserving textual capabilities with a low training cost. Notably, we achieve the best open-source result on the MMAU-speech and MMSU benchmarks and rank third among all the models.
Cross-attention is widely used in speech-to-text (S2T) systems and often exploited for downstream applications such as timestamp prediction and speech-text alignment, under the assumption that it reflects input-output dependencies. While extensively debated in NLP, its explanatory role remains underexplored in the speech domain. We empirically assess the explanatory power of cross-attention in S2T models by comparing attention scores with input saliency maps from feature-attribution methods. Our analysis spans monolingual and multilingual, single-task and multi-task models at multiple scales. We find moderate alignment between attention and saliency, particularly when aggregating across heads and layers. However, cross-attention captures only about 50% of input relevance and, at best, 52-75% of the encoder saliency. These results show that cross-attention offers useful but incomplete explanatory cues and should be interpreted with caution as a proxy for model behavior in S2T systems.
We introduce Ada-Mic, a method for adaptive close-to-mic speech detection that allows for flexible orientation ranges of the smartphone. Our method uses Generalized Cross-Correlation features as an auxiliary spatial signal that implicitly encodes the Direction-of-Arrival of speech signals to the smartphone, disentangling features for distance from those relating to orientation. It is a lightweight module that can be easily integrated into existing close-to-mic speech detectors, adding little to negligible amounts of extra computation. Our results show up to 24% improvement in accuracy over prior works for challenging distances and orientations, and demonstrates strong robustness to background noise. Ada-Mic advances the goal of No-Hot-Word wakening, enabling more natural interactions with users.
How can we adapt acoustic scene classification (ASC) models to an unlabeled target domain in the absence of original source data? This challenge is critical in real-world ASC, where device heterogeneity (e.g., varying microphone frequency responses) severely degrades adaptation performance. While Multi-Source-Free Domain Adaptation (MSFDA) offers a privacy-compliant solution, effective aggregation is hindered by inconsistent prediction biases and unknown acoustic similarities between source and target devices. We propose FᴀSᴏLᴀ, a robust MSFDA framework designed to mitigate device heterogeneity by integrating posterior adjustment to correct bias through alignment with the estimated target priors, and label agreement to prioritize reliable models based on prediction consistency. Experimental results indicate that FᴀSᴏLᴀ effectively handles multi-domain shifts and significantly improves adaptation performance.
Individualized head-related impulse responses (HRIRs) enable binaural rendering, but dense per-listener measurements are costly. We address HRIR spatial up-sampling from sparse per-listener measurements: given a few measured HRIRs for a listener, predict HRIRs at unmeasured target directions. Prior learning methods often work in the frequency domain, rely on minimum-phase assumptions or separate timing models, and use a fixed direction grid, which can degrade temporal fidelity and spatial continuity. We propose HRIR-Former, a time-domain, grid-free binaural Transformer for reconstructing HRIRs at arbitrary directions from sparse inputs. It uses sinusoidal spatial features, a Conv1D refinement module, and auxiliary interaural time difference (ITD) and interaural level difference (ILD) heads. On SONICOM, it improves normalized mean squared error (NMSE), cosine distance, and ITD/ILD errors over prior methods; ablations validate modules and show minimum-phase preprocessing is unnecessary.
High-fidelity immersive audio experiences demand personalized head-related transfer functions (HRTFs), because generic HRTFs yield inaccurate perceptual localization. Neural fields (NFs) that take the sound source direction together with a compact set of subject-specific parameters to generate personalized HRTFs have achieved accurate HRTF spatial upsampling. These parameters are typically optimized from measured HRTFs, which requires an intricate measurement process in an anechoic chamber. Towards more accessible HRTF personalization, we propose a sim-to-real NF (S2RNF) that converts HRTFs simulated from an individual's 3D head mesh into their realistic counterparts. Specifically, we infer S2RNF's subject-specific parameters from the simulated HRTFs, and use those parameters to predict the realistic HRTFs. Our experiments confirm that S2RNF outperforms the original simulations and existing sim-to-real methods.