English

Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER

Audio and Speech Processing 2026-01-30 v1 Sound

Abstract

While Automatic Speech Recognition (ASR) is typically benchmarked by word error rate (WER), real-world applications ultimately hinge on semantic fidelity. This mismatch is particularly problematic for dysarthric speech, where articulatory imprecision and disfluencies can cause severe semantic distortions. To bridge this gap, we introduce a Large Language Model (LLM)-based agent for post-ASR correction: a Judge-Editor over the top-k ASR hypotheses that keeps high-confidence spans, rewrites uncertain segments, and operates in both zero-shot and fine-tuned modes. In parallel, we release SAP-Hypo5, the largest benchmark for dysarthric speech correction, to enable reproducibility and future exploration. Under multi-perspective evaluation, our agent achieves a 14.51% WER reduction alongside substantial semantic gains, including a +7.59 pp improvement in MENLI and +7.66 pp in Slot Micro F1 on challenging samples. Our analysis further reveals that WER is highly sensitive to domain shift, whereas semantic metrics correlate more closely with downstream task performance.

Keywords

Cite

@article{arxiv.2601.21347,
  title  = {Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER},
  author = {Xiuwen Zheng and Sixun Dong and Bornali Phukon and Mark Hasegawa-Johnson and Chang D. Yoo},
  journal= {arXiv preprint arXiv:2601.21347},
  year   = {2026}
}

Comments

Accepted to ICASSP 2026

R2 v1 2026-07-01T09:25:09.670Z