English
Related papers

Related papers: On-the-fly Routing for Zero-shot MoE Speaker Adapt…

200 papers

Accurate recognition of dysarthric and elderly speech remain challenging tasks to date. Speaker-level heterogeneity attributed to accent or gender, when aggregated with age and speech impairment, create large diversity among these speakers.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Mengzhe Geng , Xurong Xie , Rongfeng Su , Jianwei Yu , Zengrui Jin , Tianzi Wang , Shujie Hu , Zi Ye , Helen Meng , Xunying Liu

Data-intensive fine-tuning of speech foundation models (SFMs) to scarce and diverse dysarthric and elderly speech leads to data bias and poor generalization to unseen speakers. This paper proposes novel structured speaker-deficiency…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Shujie Hu , Xurong Xie , Mengzhe Geng , Jiajun Deng , Zengrui Jin , Tianzi Wang , Mingyu Cui , Guinan Li , Zhaoqing Li , Helen Meng , Xunying Liu

Personalizing dysarthric ASR is hindered by demanding enrollment collection and per-user training. We propose a hybrid meta-training method for a single model, enabling zero-shot and few-shot on-the-fly personalization via in-context…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-24 Dhruuv Agarwal , Harry Zhang , Yang Yu , Quan Wang

Dysarthric speech recognition is a challenging task as dysarthric data is limited and its acoustics deviate significantly from normal speech. Model-based speaker adaptation is a promising method by using the limited dysarthric speech to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-04 Disong Wang , Jianwei Yu , Xixin Wu , Lifa Sun , Xunying Liu , Helen Meng

The application of data-intensive automatic speech recognition (ASR) technologies to dysarthric and elderly adult speech is confronted by their mismatch against healthy and nonaged voices, data scarcity and large speaker-level variability.…

Voice-based human-machine interaction is a primary modality for accessing intelligent systems, yet individuals with dysarthria face systematic exclusion due to recognition performance gaps. Whilst automatic speech recognition (ASR) achieves…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-22 Ali Alsayegh , Tariq Masood

A key challenge in dysarthric speech recognition is the speaker-level diversity attributed to both speaker-identity associated factors such as gender, and speech impairment severity. Most prior researches on addressing this issue focused on…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Mengzhe Geng , Zengrui Jin , Tianzi Wang , Shujie Hu , Jiajun Deng , Mingyu Cui , Guinan Li , Jianwei Yu , Xurong Xie , Xunying Liu

Dysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to new dysarthric speakers. Traditional speaker adaptation…

Sound · Computer Science 2024-09-25 Shiyao Wang , Shiwan Zhao , Jiaming Zhou , Aobo Kong , Yong Qin

Despite the rapid progress of automatic speech recognition (ASR) technologies targeting normal speech in recent decades, accurate recognition of dysarthric and elderly speech remains highly challenging tasks to date. Sources of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-18 Mengzhe Geng , Xurong Xie , Zi Ye , Tianzi Wang , Guinan Li , Shujie Hu , Xunying Liu , Helen Meng

In this work, we propose a zero-shot voice conversion method using speech representations trained with self-supervised learning. First, we develop a multi-task model to decompose a speech utterance into features such as linguistic content,…

Sound · Computer Science 2023-02-17 Shehzeen Hussain , Paarth Neekhara , Jocelyn Huang , Jason Li , Boris Ginsburg

Dysarthric speech recognition is a challenging task due to acoustic variability and limited amount of available data. Diverse conditions of dysarthric speakers account for the acoustic variability, which make the variability difficult to be…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Xurong Xie , Rukiye Ruzi , Xunying Liu , Lan Wang

Recently, Mixture of Experts (MoE) based Transformer has shown promising results in many domains. This is largely due to the following advantages of this architecture: firstly, MoE based Transformer can increase model capacity without…

Sound · Computer Science 2021-05-10 Zhao You , Shulin Feng , Dan Su , Dong Yu

Transformer models have been used in automatic speech recognition (ASR) successfully and yields state-of-the-art results. However, its performance is still affected by speaker mismatch between training and test data. Further finetuning a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-19 Yingzhu Zhao , Chongjia Ni , Cheung-Chi Leung , Shafiq Joty , Eng Siong Chng , Bin Ma

Dysarthric speech recognition has posed major challenges due to lack of training data and heavy mismatch in speaker characteristics. Recent ASR systems have benefited from readily available pretrained models such as wav2vec2 to improve the…

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to…

Dysarthria is a speech disorder that hinders communication due to difficulties in articulating words. Detection of dysarthria is important for several reasons as it can be used to develop a treatment plan and help improve a person's quality…

Automatic recognition of disordered speech remains a highly challenging task to date. Sources of variability commonly found in normal speech including accent, age or gender, when further compounded with the underlying causes of speech…

Sound · Computer Science 2022-01-20 Mengzhe Geng , Shansong Liu , Jianwei Yu , Xurong Xie , Shoukang Hu , Zi Ye , Zengrui Jin , Xunying Liu , Helen Meng

Mixture-of-Experts (MoE) models achieve efficient scaling through sparse expert activation, but often suffer from suboptimal routing decisions due to distribution shifts in deployment. While existing test-time adaptation methods could…

Computation and Language · Computer Science 2025-10-17 Guinan Su , Yanwu Yang , Li Shen , Lu Yin , Shiwei Liu , Jonas Geiping

A key task for speech recognition systems is to reduce the mismatch between training and evaluation data that is often attributable to speaker differences. Speaker adaptation techniques play a vital role to reduce the mismatch. Model-based…

Sound · Computer Science 2024-06-17 Xurong Xie , Xunying Liu , Tan Lee , Lan Wang

State-of-the-art automatic speech recognition (ASR) models like Whisper, perform poorly on atypical speech, such as that produced by individuals with dysarthria. Past works for atypical speech have mostly investigated fully personalized (or…

Sound · Computer Science 2025-09-23 Vishnu Raja , Adithya V Ganesan , Anand Syamkumar , Ritwik Banerjee , H Andrew Schwartz
‹ Prev 1 2 3 10 Next ›