Total: 1
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER from style mismatch), and unreliable word-level timing. We show models already encode both styles, the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot from English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We introduce supervised cross-attention finetuning improving word-level timestamps on disfluent speech beyond forced alignment baselines. Finally we introduce a new task: verbatimize, enabling scalable creation/enrichment of speech corpora with high quality canonical verbatim transcriptions.