wagner26@interspeech_2026@ISCA

Total: 1

#1 Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing [PDF] [Copy] [Kimi] [REL]

Author: Laurin Wagner

Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER from style mismatch), and unreliable word-level timing. We show models already encode both styles, the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot from English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We introduce supervised cross-attention finetuning improving word-level timestamps on disfluent speech beyond forced alignment baselines. Finally we introduce a new task: verbatimize, enabling scalable creation/enrichment of speech corpora with high quality canonical verbatim transcriptions.

Subject: INTERSPEECH.2026 - Speech Recognition