INTERSPEECH.2026 - Speech Detection

| Total: 101

#1 Beyond Short Segments : Expanding Speaker Embeddings with Vector Archives [PDF] [Copy] [Kimi] [REL]

Authors: Hyunku Kang, Minkyu Cho, Chanwoo Kim

The performance of state-of-the-art speaker verification (SV) systems severely degrades on short utterances due to insufficient speaker-specific information. To address this critical challenge, we propose the Vector Archive Mapping ECAPA (VAM-ECAPA), a novel system designed to enhance feature extraction from short-duration speech. The core of our system is the Transformer-based Vector Archive Mapping with Statistical Pooling (TVAMSP) module, which enriches information-scarce features by mapping them against a learnable Vector Archive of canonical speaker traits. By integrating the TVAMSP module into a strong WavLM+ECAPA-TDNN baseline, our system learns to map sparse features from short segments into robust, discriminative speaker representations. Experiments on the VoxCeleb1 benchmark show that our proposed VAM-ECAPA achieves a highly competitive EER of 8.334% on 1-second test segments, a 54.8% relative error reduction compared to a conventionally-trained baseline.

Subject: INTERSPEECH.2026 - Speech Detection


#2 Revisiting Label-Free Speaker Embedding Enhancement with vMF Profile Likelihood [PDF] [Copy] [Kimi] [REL]

Authors: Seunghwan Kim, Jinyong Kim, Sooyoung Yang, Youngjin Ko, Myungjoo Kang

Embedding enhancement improves speaker verification under acoustic mismatch without modifying a frozen backbone. Recent work has established a practical label-free setting for this task, but often adopts increasingly structured formulations. Here, the clean target is directly observed during training, making enhancement a matching problem on the unit hypersphere. We model the clean target with a von Mises-Fisher (vMF) likelihood and profile out a sample-wise concentration parameter, yielding a simple closed-form objective with adaptive weighting. Across VoxCeleb1, VoxSRC23, CN-Celeb, VOiCES, and VC-Mix, the proposed method largely preserves the baseline and gives clearer gains on challenging mismatch sets. It also remains stable under a broad single-view recipe, where a recent diffusion baseline becomes less reliable in controlled comparisons. These results suggest that effective label-free embedding enhancement in this setting does not require a highly structured formulation.

Subject: INTERSPEECH.2026 - Speech Detection


#3 On the Robustness of Speaker Embeddings for Cross-Domain Speaker Retrieval [PDF] [Copy] [Kimi] [REL]

Authors: Chuanqi Huang, Wei Xie, Xilu Wang

Deploying speaker retrieval systems requires robust cross-domain embedding generalization. However, existing benchmarks focus on verification metrics, leaving ranking stability under retrieval constraints under-explored. This paper evaluates six pre-trained embedding models across multiple cross-domain scenarios. First, while supervised multi-scale models resist channel filtering and aging drift, most architectures overfit to language-specific phonetic variations under cross-lingual mismatch. Second, we leverage adaptive symmetric normalization as a training-free backend strategy to improve retrieval performance. By selecting high-scoring background cohorts to estimate localized score statistics, this strategy normalizes global shifts caused by channel distortions and restores ranking consistency across all models. These insights demonstrate that combining multi-scale local features with adaptive backend calibration is effective for cross-domain speaker retrieval.

Subject: INTERSPEECH.2026 - Speech Detection


#4 Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings [PDF] [Copy] [Kimi] [REL]

Authors: Byoungjun So, Jaejun Lee, Kyogu Lee

Understanding speaker attributes is crucial for voice-related applications, yet conventional approaches rely on fixed categorical labels, lacking semantic richness and zero-shot generalizability. We propose a novel framework for open-set speaker attribute prediction leveraging Large Language Model (LLM) embed-dings to represent attributes in a continuous semantic space. To bridge the cross-modal gap, we introduce a keyword-appending strategy that structures broad semantic representations into a compact, discriminative manifold. Furthermore, we employ a top-k negative loss to establish robust decision boundaries in crowded semantic regions. Experimental results on LibriTTS-P demonstrate that our method outperforms closed-set benchmarks and generalizes effectively to unseen synonyms. Geometric analysis suggests that our strategies regularize the embedding manifold, balancing semantic cohesion with predictive clarity.

Subject: INTERSPEECH.2026 - Speech Detection


#5 Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity [PDF] [Copy] [Kimi] [REL]

Authors: Ismail Rasim Ulgen, John Hansen, Carlos Busso, Berrak Sisman

Modeling speech variation is key to natural, expressive generation. Speaker embeddings are commonly used to condition personalized speech systems, but they are typically trained for speaker recognition, where intra-speaker variability is suppressed and inter-speaker separation is maximized. This objective leads to overly compact representations that may discard variations crucial for generation. We revisit this design choice and propose a sub-center modeling framework for speaker embeddings. Instead of a single prototype per speaker, we learn multiple sub-centers during discriminative training, allowing utterances to align with different prototypes. This strategy preserves structured intra-speaker variability while maintaining discriminability. In zero-shot voice conversion, our method improves intelligibility, increases pitch variability, achieves higher naturalness ratings, and retains strong speaker verification performance.

Subject: INTERSPEECH.2026 - Speech Detection


#6 NoiseLoRA-SV: Hierarchical Noise-Conditioned Adaptation with Embedding Distillation for Robust Speaker Verification [PDF] [Copy] [Kimi] [REL]

Authors: Dai Gao, Chen Jiang, Sizhe Liu, Peng Zhang

Current speaker verification (SV) models, including Low-Rank Adaptation (LoRA) variants, rely on static inference-time parameters and show limited robustness to non-stationary noise. We propose NoiseLoRA-SV, a dynamic framework that generates instance-adaptive weights on-the-fly during inference. Instead of full end-to-end fine-tuning, it injects noise-conditioned LoRA modules into a lightly fine-tuned backbone with moderate additional parameter overhead. A Convolutional Recurrent Network (CRN) extracts hierarchical noise representations: global embeddings drive a hypernetwork to generate LoRA matrices, while local embeddings control a frame-level time-varying gate. Optimized via InfoNCE-based contrastive distillation, NoiseLoRA-SV aligns noisy embeddings with clean speaker representations. Evaluated on the VoxCeleb1 corpus with MUSAN and unseen NonSpeech100 noises, NoiseLoRA-SV consistently yields lower equal error rates (EERs) than static baselines.

Subject: INTERSPEECH.2026 - Speech Detection


#7 Comparing Self-Supervised and Domain-Invariant Features for Cross-Domain Voice Phishing Detection [PDF] [Copy] [Kimi] [REL]

Authors: Jeongmin Lee, Seung Yun, Minkyu Lee, Ran Han, Yoonkyu Woo, Jinxia Huang

Voice phishing detection faces three critical challenges: real criminal recordings are unavailable due to privacy constraints; when available, only a handful of samples exist, insufficient for fine-tuning; and lightweight acoustic-only detection is needed as an alternative to large self-supervised models. We compare domain-invariant prosodic features and self-supervised representations (HuBERT, wav2vec2.0) through cross-domain evaluation--training on scenario-based actor recordings and testing on authentic criminal calls. Domain-invariant prosodic features achieve 69.5% F1 zero-shot and 71.0% with 5-shot learning. HuBERT achieves highest performance (94.2% F1, 5-shot), while wav2vec2.0 exhibits a precision-oriented detection profile (90.2% F1 with 99.4% precision, 5-shot). These findings reveal fundamental trade-offs: domain-invariant features enable zero-shot deployment when no real data exists, while SSL methods achieve higher performance but require real samples and compute.

Subject: INTERSPEECH.2026 - Speech Detection


#8 Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis [PDF] [Copy] [Kimi] [REL]

Authors: Zhuolin Yi, Jun Xue, Yanzhen Ren, Yihuan Huang, Yi Chai, Daixian Li, Guanxiang Feng, Jiajun Liu

The scale of speech anti-spoofing datasets has grown exponentially over the past decade, driven by the assumption that larger data leads to better performance. However, it remains unclear whether indiscriminate scaling commensurately improves model generalization. This study challenges the "scale-first" paradigm by decoupling the impacts of training data scale versus diversity. Through experiments on representative datasets, we report two key findings: (1) Larger is not always better. Expanding data scale excessively under fixed generation methods yields negligible returns and may even degrade cross-domain generalization due to overfitting.(2) Diversity outweighs scale. A smaller composite training set featuring diverse attacks significantly outperforms larger-scale datasets with limited diversity in cross-dataset evaluations. We conclude that future dataset construction should prioritize the diversity of generation methods over scale to effectively enhance model generalization.

Subject: INTERSPEECH.2026 - Speech Detection


#9 ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks [PDF] [Copy] [Kimi] [REL]

Authors: Aurosweta Mahapatra, Ismail Rasim Ulgen, Kong Aik Lee, Nicholas Andrews, Berrak Sisman

Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific artifacts rather than transferable cues of natural speech. In contrast, humans internalize variability in real speech and detect fakes as deviations from it. We introduce ProSDD, a two-stage framework that enriches model embeddings through supervised masked prediction of speaker-conditioned prosodic variation based on pitch, voice activity, and energy. Stage I learns prosodic variability from real speech, and Stage II jointly optimizes this objective with spoof classification. ProSDD consistently outperforms baselines under both ASVspoof 2019 and 2024 training, reducing ASVspoof 2024 EER from 25.43% to 16.14% (2019-trained) and from 39.62% to 7.38% (2024-trained), while achieving 50% relative reductions on EmoFake and EmoSpoof-TTS.

Subject: INTERSPEECH.2026 - Speech Detection


#10 Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack [PDF] [Copy] [Kimi] [REL]

Authors: Nitin Choudhury, Bikrant Bikram Pratap Maurya, Arun Balaji Buduru, Orchid Chetia Phukan

Acoustic side-channel attacks (ASCA) on keyboards have gained increasing attention, yet impact of speech representation learning models in ASCA remains unexplored. Addressing this, we introduce KEYAC, a dataset designed to analyze representation generalization for ASCA under both standard and VoIP codec settings. On KEYAC, we evaluate six representation learning models under zero-shot and partial fine-tuning settings using fully connected and convolutional networks. Results show that while partial fine-tuning improves performance, models struggle to generalize across VoIP codecs. We hypothesize this limitation stems from inadequate modeling of nonlinear feature interactions in conventional fine-tuning architectures. To address this, we employ Kolmogorov-Arnold Networks (KAN) for fine-tuning. Empirical results show that KAN-based fine-tuning consistently outperforms the baselines and establishes a new state-of-the-art on KEYAC.

Subject: INTERSPEECH.2026 - Speech Detection


#11 Ouroboros: Self-Referential Backdoor Attacks on Speech Enhancement via Clean Audio Triggers [PDF] [Copy] [Kimi] [REL]

Authors: Yunjie Zhou, Yuheng Huang, Diqun Yan

Speech enhancement models are widely deployed as front-end modules in real-time speech services, yet their vulnerability to backdoor attacks remains unexplored. Existing backdoor methods are confined to classification tasks and rely on active trigger injection, an assumption incompatible with the passive processing nature of speech enhancement models. In this paper, we propose Ouroboros, a novel backdoor attack framework that leverages the ideal clean outputs of speech enhancement models as natural triggers, enabling inference-time activation without any external trigger injection. Extensive evaluations show Ouroboros achieves near-perfect attack success rates with minimal performance degradation on diverse models and datasets. Physical-world validations confirm that naturally recorded, unaltered clean audio can reliably activate the backdoor. Moreover, Ouroboros generalizes to targeted content-tampering attacks and remains effective against common filtering and fine-tuning defenses.

Subject: INTERSPEECH.2026 - Speech Detection


#12 FreqGuard: Leveraging Frequency-Domain Feature Priors for Universal Proactive Voice Defense [PDF] [Copy] [Kimi] [REL]

Authors: Yankai Wang, Zhipeng Chen, Yuxuan Du, Rong Zheng, Jing Deng

The rapid development of voice deepfakes poses significant risks to privacy and security. This paper introduces FreqGuard, a universal proactive defense framework designed to protect against malicious speech synthesis attacks. FreqGuard leverages learnable frequency-domain feature priors to generate imperceptible perturbations that effectively disrupt voice synthesis systems. Unlike methods based on random noise, FreqGuard explores frequency operations to obtain priors that degrade speaker embeddings while minimizing perceptual distortion. These priors guide model training to target the shared subspace between verification and synthesized features. By jointly optimizing multiple losses, FreqGuard achieves high imperceptibility and low speaker similarity. Experimental results show that FreqGuard reduces attack success rates against black-box TTS systems while preserving speech quality, achieving favorable cross-model generalization and perceptual quality compared with existing methods.

Subject: INTERSPEECH.2026 - Speech Detection


#13 SEA-Spoof: Bridging the Gap in Multilingual Audio Deepfake Detection for South-East Asia [PDF] [Copy] [Kimi] [REL]

Authors: Jinyang Wu, Nana Hou, Zihan Pan, Qiquan Zhang, Sailor Hardik, Soumik Mondal

The rapid growth of the digital economy in South-East Asia (SEA) has amplified the risks of audio deepfakes, yet existing datasets provide limited coverage of SEA languages, hindering robust detection. We present SEA-Spoof, the first large-scale audio deepfake detection dataset dedicated to six SEA languages: Tamil, Hindi, Thai, Indonesian, Malay, and Vietnamese. SEA-Spoof contains over 700 hours of paired real and spoof speech generated by diverse state-of-the-art open-source and closed-source systems. Its balanced, transcript aligned design enables controlled language and system level evaluation. Benchmarking reveals severe cross-lingual degradation of models trained on high resource languages, while fine-tuning on SEA-Spoof restores performance across languages and synthesis sources. SEA-Spoof establishes a foundation for robust, cross-lingual and region-aware deepfake detection in SEA.

Subject: INTERSPEECH.2026 - Speech Detection


#14 When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus [PDF] [Copy] [Kimi] [REL]

Authors: Kirill Borodin, Vasiliy Kudryavtsev, Maxim Maslov, Mikhail Gorodnichev, Grach Mkrtchian

We introduce LRLspoof, a large-scale multilingual synthetic-speech corpus for cross-lingual spoof detection, comprising 2,732 hours of audio generated with 24 open-source TTS systems across 66 languages, including 45 low-resource languages under our operational definition. To evaluate robustness without requiring target-domain bona fide speech, we benchmark 11 publicly available countermeasures using threshold transfer: for each model we calibrate an EER operating point on pooled external benchmarks and apply the resulting threshold, reporting spoof rejection rate (SRR). Results show model-dependent cross-lingual disparity, with spoof rejection varying markedly across languages even under controlled conditions, highlighting language as an independent source of domain shift in spoof detection. The dataset is publicly available at https://huggingface.co/datasets/lab260/LRLspoof and https://modelscope.cn/datasets/lab260/LRLspoof.

Subject: INTERSPEECH.2026 - Speech Detection


#15 MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection [PDF] [Copy] [Kimi] [REL]

Authors: Xueping Zhang, Zhenshan Zhang, Yechen Wang, Linxi Li, Liwei Jin, Ming Li

Existing speech anti-spoofing benchmarks rely on a narrow set of public models, creating a substantial gap from real-world scenarios in which commercial systems employ diverse, often proprietary APIs. To address this issue, we introduce Multi-API Spoof, a multi-API audio anti-spoofing dataset comprising about 230 hours of synthetic speech generated by 30 distinct APIs, including commercial services, open-source models, and online platforms. Furthermore, we propose Nes2Net-LA, a local-attention enhanced variant of Nes2Net that improves local context modeling and fine-grained spoofing feature extraction. Based on this dataset, we also define the API tracing task, enabling fine-grained attribution of spoofed audio to its generation source. Experiments show that Nes2Net-LA achieves state-of-the-art performance and offers superior robustness, particularly under diverse and unseen spoofing conditions. Code 1 and dataset 2 have been released.

Subject: INTERSPEECH.2026 - Speech Detection


#16 Aleatoric Style Uncertainty Augmentation with GMM for Domain Generalization in Anti-spoofing [PDF] [Copy] [Kimi] [REL]

Authors: Jin Li, Man-Wai Mak, Johan Rohdin, Oldřich Plchot, Kong Aik Lee, Bo Wen, Yunfeng Liu

Speech anti-spoofing systems often degrade in out-of-domain (OOD) settings due to varied unknown spoofing attacks. Domain generalization (DG) methods address this by improving model robustness across diverse domains. Style augmentation is a DG approach that synthesizes new features by modeling domain-shift uncertainty using feature statistics learned during training. However, previous style augmentation methods rely on a single batch-level uncertainty by assuming a unimodal style distribution, which may not hold in mixed domain batches. To address this issue, we propose Gaussian mixture model (GMM) based aleatoric style uncertainty (ASU) to model within-component variability. In addition, we proposed an online EM-like update for the GMM in an end-to-end way. Extensive experiments on anti-spoofing and spoofing-aware speaker verification (SASV) show that ASU significantly outperforms existing methods and surpasses state-of-the-art systems. Code is available at GitHub.

Subject: INTERSPEECH.2026 - Speech Detection


#17 Diffusion Reconstruction towards Generalizable Audio Deepfake Detection [PDF] [Copy] [Kimi] [REL]

Authors: Bo Cheng, Songjun Cao, Xiaoming Zhang, Jie Chen, Long Ma, Fei Chen

Achieving robust generalization against unseen attacks remains a challenge in Audio Deepfake Detection (ADD), driven by the rapid evolution of generative models. To address this, we propose a framework centered on hard sample classification. The core idea is that a model capable of distinguishing challenging hard samples is inherently equipped to handle simpler cases effectively. We investigate multiple reconstruction paradigms, identifying the diffusion-based method as optimal for generating hard samples. Furthermore, we leverage multi-layer feature aggregation and introduce a Regularization-Assisted Contrastive Learning (RACL) objective to enhance generalizability. Experiments demonstrate the superior generalization of our approach, with our best model achieving a significant reduction in the average Equal Error Rate (EER) compared to the baseline.

Subject: INTERSPEECH.2026 - Speech Detection


#18 Hard Positive-targeted Training for Robust Audio Deepfake Detection under Neural Codec Processing [PDF] [Copy] [Kimi] [REL]

Authors: Jiwon Seo, Inho Kim, Seongkyu Han, Thien-Phuc Doan, Souhwan Jung

Neural audio codecs are increasingly used in speech pipelines, enabling high-quality compression at low bitrates. However, neural codec (NC) processing may impair audio deepfake detection (ADD) by distorting discriminative cues and introducing artifacts that may obscure spoofing traces. Our embedding analysis indicates that the robustness drop is driven mainly by bonafide-side errors: NC-processed bonafide speech shifts toward the spoof region more than NC-processed spoof shifts toward bonafide. To mitigate this effect, we propose a training strategy that (i) introduces an auxiliary loss to regularize frequently misclassified NC-processed bonafide samples and (ii) uses a mini-batch construction scheme that repeatedly presents boundary-adjacent NC-processed bonafide-spoof pairs during optimization. Experiments using equal error rate (EER) and accuracy demonstrate improved robustness under NC conditions while maintaining spoof detection performance.

Subject: INTERSPEECH.2026 - Speech Detection


#19 Deepfake Word Detection by Next-token Prediction using Fine-tuned Whisper [PDF] [Copy] [Kimi] [REL]

Authors: Hoan My Tran, Xin Wang, Wanying Ge, Xuechen Liu, Junichi Yamagishi

Deepfake speech utterances can be forged by replacing one or more words in a bona fide utterance with semantically different words synthesized with speech-generative models. While a dedicated synthetic word detector could be developed, we developed a cost-effective method that fine-tunes a pre-trained Whisper model to detect synthetic words while transcribing the input utterance via next-token prediction. We further investigate using partially vocoded utterances as the fine-tuning data, thus reducing the cost of data collection. Our experiments demonstrate that, on in-domain test data, the fine-tuned Whisper yields low synthetic-word detection error rates and transcription error rates. On out-of-domain test data with synthetic words produced with unseen speech-generative models, the fine-tuned Whisper remains on par with a dedicated ResNet-based detection model; however, the overall performance degradation calls for strategies to improve its generalization capability.

Subject: INTERSPEECH.2026 - Speech Detection


#20 Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing [PDF] [Copy] [Kimi] [REL]

Authors: Awais Khan, Kutub Uddin, Khalid Malik

Attributing a synthetic utterance to its originating system remains an open challenge: closed-set models fail to reject unseen synthesizers and produce overconfident predictions. To address this, we propose a dual-branch gated fusion framework that pairs XLSR-53 with CORES, a 66-dimensional descriptor that, unlike prior Linear Filter Bank-only work, spans cepstral, oscillatory, rhythmic, energy, and spectral dimensions to capture complementary synthesis artifacts. Our analysis shows XLSR-53 remains discriminative in-domain (ID) while CORES generalizes stably under distribution shift (OOD), yet their naive concatenation fails due to SSL representational imbalance. To resolve this, an input-conditioned gate adaptively weights each branch under joint training with cross-entropy, an energy margin loss for ID/OOD separation, and a gate diversity term. On the MLAAD benchmark, our system achieves 97.6% ID accuracy, 4.9% EERc, and an 83.5% relative reduction in FPR95 over the 2025 baseline.

Subject: INTERSPEECH.2026 - Speech Detection


#21 Mixture of Spectral Experts for Audio Deepfake Detection [PDF] [Copy] [Kimi] [REL]

Authors: Yaxuan Qiu, Zhe Li, Mieradilijiang Maimaiti, Zunwang Ke, Wushour Silamu

Recent advances in neural speech synthesis have produced highly natural waveforms, making audio deepfake detection increasingly challenging as spoofing artifacts become less perceptible. Although pre-trained speech models provide robust representations, they may overlook low-level physical cues, particularly magnitude and phase information. To address this limitation, we propose a detection framework that combines a frequency audio encoder (FAE) with spectral parameter-efficient fine-tuning. The FAE explicitly models magnitude and phase cues, while the proposed Mixture of Spectral Experts (MoSE) efficiently adapts the pre-trained speech model to generation-dependent distribution shifts. By applying low-rank updates in the singular value decomposition (SVD) domain while keeping the singular bases frozen, MoSE facilitates task-specific adaptation to spoofing-related spectral artifacts. Evaluations on ASVspoof 2019 LA, ASVspoof 2021 LA/DF, and In-the-Wild benchmarks demonstrate the effectiveness of our approach and its strong generalization to unseen channel variations and real-world spoofing attacks.

Subject: INTERSPEECH.2026 - Speech Detection


#22 Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection [PDF] [Copy] [Kimi] [REL]

Authors: Zhuodong Liu, Hugen Lv, Xiangyu Li, Chunhong Yuan

Audio deepfake detectors often fail to generalize across speakers, as they learn speaker-identity features rather than synthesis artifacts, known as implicit identity leakage. Existing methods address this but incur architectural complexity or training instability. This paper proposes a dual-granularity orthogonal disentanglement framework enforcing feature independence at two levels: sample-level cosine orthogonality captures directional decorrelation, while batch-level cross-covariance regularization eliminates linear correlations across embedding dimensions. A curriculum disentanglement schedule progressively strengthens the orthogonality constraint without auxiliary networks or adversarial dynamics. Experiments on ASVspoof 2019 LA, ASVspoof 2021 DF, and In-the-Wild datasets demonstrate that the proposed method achieves 1.35%, 7.88%, and 21.58% equal error rates (EER), respectively, surpassing gradient reversal disentanglement by 2.60% absolute on cross-dataset transfer.

Subject: INTERSPEECH.2026 - Speech Detection


#23 RAT: Reference-Augmented Training for ASV Anti-Spoofing [PDF] [Copy] [Kimi] [REL]

Authors: Vojtěch Staněk, Anton Firc, Jakub Reš, Kamil Malinka

We introduce a spoofing countermeasure architecture conditioned on speaker-reference recordings, but observe that it converges to a solution that effectively ignores the reference during inference. Surprisingly, training with a reference channel induces invariance that improves deepfake detection, even when the reference is absent or mismatched during inference. Based on this observation, we propose a Reference-Augmented Training (RAT) strategy. RAT yields improved detection performance compared to single-utterance baselines, even when the reference recording is replaced with a zero vector at inference. Through rigorous analysis, we demonstrate that the optimization process rapidly diminishes the reference contributions, leading to inference largely independent of the reference channel. Using RAT, we achieve state-of-the-art 2.57% EER and 0.074 minDCF on the ASVspoof 5 benchmark with a single detector, surpassing even large ensemble systems.

Subject: INTERSPEECH.2026 - Speech Detection


#24 SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing [PDF] [Copy] [Kimi] [REL]

Authors: Anton Firc, Vojtěch Staněk, Zbyněk Lička, Kamil Malinka, Martin Perešíni

We present SpAArSIST, a deployment-oriented refinement of the widely used AASIST graph pooling backend for self-supervised learning (SSL) based anti-spoofing. Motivated by redundant operations in public implementations, we replace learned pooling and stack-node attention with explicit, lightweight choices: separate train and inference graph pooling ratios (ktr, kinf), magnitude-based node scoring, and mean aggregation of graph nodes. The best overall configuration (rank 1) cuts backend compute by 20.7% (195.045M → 154.706M MACs) and model size by 4.1% (611.8k → 586.4k params), while improving out-of-domain robustness on In-the-Wild to 2.82% EER and 0.078 minDCF (from 4.64% and 0.133) and remaining competitive on ASVspoof 5. We further provide a composite selection score that summarizes accuracy, calibration, and compute to support balanced deployment-oriented model choice.

Subject: INTERSPEECH.2026 - Speech Detection


#25 Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck [PDF] [Copy] [Kimi] [REL]

Authors: Anh-Tuan DAO, Driss Matrouf, Mickael Rouvier, Nicholas Evans

Rapid advancements in generative speech technology have compromised the reliability of voice biometrics. While current spoofing detectors excel when assessed under in-domain conditions, generalisation to out-of-domain settings is often poor. We show that this can be due to linguistic bias. A reliance on linguistic cues observed in training data can then compromise robustness to cross-data. We propose a linguistic-invariant spoofing detection framework utilizing teacher-student adver-sarial learning. The linguistic-aware teacher model, pre-trained on linguistic content of an external dataset, guides the student detector via gradient reversal to minimize the linguistic information. To prevent the inadvertent removal of non-linguistic cues, we incorporate a Variational Information Bottleneck to enable suppression of principal cues. Across nine DF Arena datasets, our method achieves up to a 36.2% relative reduction in the EER compare to the baseline.

Subject: INTERSPEECH.2026 - Speech Detection