English

Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database

Sound 2025-08-27 v1 Artificial Intelligence

Abstract

Dysarthric speech recognition faces challenges from severity variations and disparities relative to normal speech. Conventional approaches individually fine-tune ASR models pre-trained on normal speech per patient to prevent feature conflicts. Counter-intuitively, experiments reveal that multi-speaker fine-tuning (simultaneously on multiple dysarthric speakers) improves recognition of individual speech patterns. This strategy enhances generalization via broader pathological feature learning, mitigates speaker-specific overfitting, reduces per-patient data dependence, and improves target-speaker accuracy - achieving up to 13.15% lower WER versus single-speaker fine-tuning.

Keywords

Cite

@article{arxiv.2508.18732,
  title  = {Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database},
  author = {Qing Xiao and Yingshan Peng and PeiPei Zhang},
  journal= {arXiv preprint arXiv:2508.18732},
  year   = {2025}
}
R2 v1 2026-07-01T05:05:54.734Z