lee26s@interspeech_2026@ISCA

Total: 1

#1 Comparing Self-Supervised and Domain-Invariant Features for Cross-Domain Voice Phishing Detection [PDF] [Copy] [Kimi] [REL]

Authors: Jeongmin Lee, Seung Yun, Minkyu Lee, Ran Han, Yoonkyu Woo, Jinxia Huang

Voice phishing detection faces three critical challenges: real criminal recordings are unavailable due to privacy constraints; when available, only a handful of samples exist, insufficient for fine-tuning; and lightweight acoustic-only detection is needed as an alternative to large self-supervised models. We compare domain-invariant prosodic features and self-supervised representations (HuBERT, wav2vec2.0) through cross-domain evaluation--training on scenario-based actor recordings and testing on authentic criminal calls. Domain-invariant prosodic features achieve 69.5% F1 zero-shot and 71.0% with 5-shot learning. HuBERT achieves highest performance (94.2% F1, 5-shot), while wav2vec2.0 exhibits a precision-oriented detection profile (90.2% F1 with 99.4% precision, 5-shot). These findings reveal fundamental trade-offs: domain-invariant features enable zero-shot deployment when no real data exists, while SSL methods achieve higher performance but require real samples and compute.

Subject: INTERSPEECH.2026 - Speech Detection