Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
Abstract
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We further introduce supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines. Finally, we propose verbatimize, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions.
Keywords
Cite
@article{arxiv.2607.18934,
title = {Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing},
author = {Laurin Wagner and Mario Zusag and Bernhard Thallinger},
journal= {arXiv preprint arXiv:2607.18934},
year = {2026}
}
Comments
Accepted at Interspeech 2026 long track