English

On-the-Fly Feature Based Rapid Speaker Adaptation for Dysarthric and Elderly Speech Recognition

Audio and Speech Processing 2023-05-30 v3 Artificial Intelligence Machine Learning Sound

Abstract

Accurate recognition of dysarthric and elderly speech remain challenging tasks to date. Speaker-level heterogeneity attributed to accent or gender, when aggregated with age and speech impairment, create large diversity among these speakers. Scarcity of speaker-level data limits the practical use of data-intensive model based speaker adaptation methods. To this end, this paper proposes two novel forms of data-efficient, feature-based on-the-fly speaker adaptation methods: variance-regularized spectral basis embedding (SVR) and spectral feature driven f-LHUC transforms. Experiments conducted on UASpeech dysarthric and DementiaBank Pitt elderly speech corpora suggest the proposed on-the-fly speaker adaptation approaches consistently outperform baseline iVector adapted hybrid DNN/TDNN and E2E Conformer systems by statistically significant WER reduction of 2.48%-2.85% absolute (7.92%-8.06% relative), and offline model based LHUC adaptation by 1.82% absolute (5.63% relative) respectively.

Keywords

Cite

@article{arxiv.2203.14593,
  title  = {On-the-Fly Feature Based Rapid Speaker Adaptation for Dysarthric and Elderly Speech Recognition},
  author = {Mengzhe Geng and Xurong Xie and Rongfeng Su and Jianwei Yu and Zengrui Jin and Tianzi Wang and Shujie Hu and Zi Ye and Helen Meng and Xunying Liu},
  journal= {arXiv preprint arXiv:2203.14593},
  year   = {2023}
}

Comments

Accepted to INTERSPEECH 2023

R2 v1 2026-06-24T10:28:03.773Z