aldeneh26@interspeech_2026@ISCA

Total: 1

#1 Which Data Matter? Embedding-Based Data Selection for Speech Recognition [PDF] [Copy] [Kimi] [REL]

Authors: Zakaria Aldeneh, Skyler Seto, Maureen de Seyssel, Jie Chi, Zijin Gu, Takuya Higuchi, Jee-weon Jung, Shinji Watanabe, David Grangier, Barry-John Theobald, Tatiana Likhomanenko

Modern ASR systems are typically trained on large-scale pseudo-labeled, in-the-wild data spanning multiple domains. While such heterogeneous data benefit generalist models designed for broad deployment, they pose challenges for specialist models targeting specific domains: specialist models lack the capacity to learn from all available data. In this work, we study targeted data selection to address these challenges, selecting relevant subsets from 100k hours of in-the-wild training data to optimize performance on target domains. We represent speech samples using embeddings that capture complementary characteristics—speaker attributes, phonetic content, and semantic meaning—and study how relevance and diversity along these axes when performing data selection affect ASR performance. Our experiments with CTC-based models show that training on a strategically selected 5% subset can exceed the performance of models trained on the full data by up to 36.8% relative WER reduction.

Subject: INTERSPEECH.2026 - Speech Recognition