huang26m@interspeech_2026@ISCA

Total: 1

#1 On the Robustness of Speaker Embeddings for Cross-Domain Speaker Retrieval [PDF] [Copy] [Kimi] [REL]

Authors: Chuanqi Huang, Wei Xie, Xilu Wang

Deploying speaker retrieval systems requires robust cross-domain embedding generalization. However, existing benchmarks focus on verification metrics, leaving ranking stability under retrieval constraints under-explored. This paper evaluates six pre-trained embedding models across multiple cross-domain scenarios. First, while supervised multi-scale models resist channel filtering and aging drift, most architectures overfit to language-specific phonetic variations under cross-lingual mismatch. Second, we leverage adaptive symmetric normalization as a training-free backend strategy to improve retrieval performance. By selecting high-scoring background cohorts to estimate localized score statistics, this strategy normalizes global shifts caused by channel distortions and restores ranking consistency across all models. These insights demonstrate that combining multi-scale local features with adaptive backend calibration is effective for cross-domain speaker retrieval.

Subject: INTERSPEECH.2026 - Speech Detection