seo26b@interspeech_2026@ISCA

Total: 1

#1 Hard Positive-targeted Training for Robust Audio Deepfake Detection under Neural Codec Processing [PDF] [Copy] [Kimi] [REL]

Authors: Jiwon Seo, Inho Kim, Seongkyu Han, Thien-Phuc Doan, Souhwan Jung

Neural audio codecs are increasingly used in speech pipelines, enabling high-quality compression at low bitrates. However, neural codec (NC) processing may impair audio deepfake detection (ADD) by distorting discriminative cues and introducing artifacts that may obscure spoofing traces. Our embedding analysis indicates that the robustness drop is driven mainly by bonafide-side errors: NC-processed bonafide speech shifts toward the spoof region more than NC-processed spoof shifts toward bonafide. To mitigate this effect, we propose a training strategy that (i) introduces an auxiliary loss to regularize frequently misclassified NC-processed bonafide samples and (ii) uses a mini-batch construction scheme that repeatedly presents boundary-adjacent NC-processed bonafide-spoof pairs during optimization. Experiments using equal error rate (EER) and accuracy demonstrate improved robustness under NC conditions while maintaining spoof detection performance.

Subject: INTERSPEECH.2026 - Speech Detection