English
Related papers

Related papers: Zero-Shot Recognition of Dysarthric Speech Using C…

200 papers

Dysarthria, a common issue among stroke patients, severely impacts speech intelligibility. Inappropriate pauses are crucial indicators in severity assessment and speech-language therapy. We propose to extend a large-scale speech recognition…

Computation and Language · Computer Science 2024-03-01 Jeehyun Lee , Yerin Choi , Tae-Jin Song , Myoung-Wan Koo

The Fearless Steps APOLLO Community Resource provides unparalleled opportunities to explore the potential of multi-speaker team communications from NASA Apollo missions. This study focuses on discovering the characteristics that make Apollo…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-19 Alkis Koudounas , Flavio Giobergia

The idea of combining multiple languages' recordings to train a single automatic speech recognition (ASR) model brings the promise of the emergence of universal speech representation. Recently, a Transformer encoder-decoder model has been…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-08 Siyuan Feng , Piotr Żelasko , Laureano Moro-Velázquez , Ali Abavisani , Mark Hasegawa-Johnson , Odette Scharenborg , Najim Dehak

Dysarthria is a disability that causes a disturbance in the human speech system and reduces the quality and intelligibility of a person's speech. Because of this effect, the normal speech processing systems can not work properly on impaired…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-22 Aref Farhadipour , Hadi Veisi

Dysfluencies and variations in speech pronunciation can severely degrade speech recognition performance, and for many individuals with moderate-to-severe speech disorders, voice operated systems do not work. Current speech recognition…

Automatic Speech Recognition (ASR) systems have been evolving quickly and reaching human parity in certain cases. The systems usually perform pretty well on reading style and clean speech, however, most of the available systems suffer from…

Computation and Language · Computer Science 2019-10-15 Quang Minh Nguyen , Thai Binh Nguyen , Ngoc Phuong Pham , The Loc Nguyen

Speech-to-text errors made by automatic speech recognition (ASR) systems negatively impact downstream models. Error correction models as a post-processing text editing method have been recently developed for refining the ASR outputs.…

Computation and Language · Computer Science 2023-06-22 Ziji Zhang , Zhehui Wang , Rajesh Kamma , Sharanya Eswaran , Narayanan Sadagopan

Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven…

Sound · Computer Science 2025-08-22 Dena Mujtaba , Nihar Mahapatra

The success of large language models (LLMs) has prompted efforts to integrate speech and audio data, aiming to create general foundation models capable of processing both textual and non-textual inputs. Recent advances, such as GPT-4o,…

Computation and Language · Computer Science 2024-10-18 Fan Bu , Yuhao Zhang , Xidong Wang , Benyou Wang , Qun Liu , Haizhou Li

Training large foundation models using self-supervised objectives on unlabeled data, followed by fine-tuning on downstream tasks, has emerged as a standard procedure. Unfortunately, the efficacy of this approach is often constrained by both…

We present a comprehensive evaluation of pretrained speech embedding systems for the detection of dysarthric speech using existing accessible data. Dysarthric speech datasets are often small and can suffer from recording biases as well as…

Dysarthric speech poses significant challenges for individuals with dysarthria, impacting their ability to communicate socially. Despite the widespread use of Automatic Speech Recognition (ASR), accurately recognizing dysarthric speech…

Sound · Computer Science 2025-02-14 Yan Wang , Mengyi Sun , Xinchen Kang , Jingting Li , Pengfei Guo , Ming Gao , Su-Jing Wang

Speech technologies are transforming interactions across various sectors, from healthcare to call centers and robots, yet their performance on African-accented conversations remains underexplored. We introduce Afrispeech-Dialog, a benchmark…

Speech has emerged as a widely embraced user interface across diverse applications. However, for individuals with dysarthria, the inherent variability in their speech poses significant challenges. This paper presents an end-to-end…

Sound · Computer Science 2024-09-17 Shuiyun Liu , Yuxiang Kong , Pengcheng Guo , Weiji Zhuang , Peng Gao , Yujun Wang , Lei Xie

Automatic recognition of disordered and elderly speech remains a highly challenging task to date due to the difficulty in collecting such data in large quantities. This paper explores a series of approaches to integrate domain adapted SSL…

Sound · Computer Science 2023-06-23 Shujie Hu , Xurong Xie , Zengrui Jin , Mengzhe Geng , Yi Wang , Mingyu Cui , Jiajun Deng , Xunying Liu , Helen Meng

As for the humanoid robots, the internal noise, which is generated by motors, fans and mechanical components when the robot is moving or shaking its body, severely degrades the performance of the speech recognition accuracy. In this paper,…

Sound · Computer Science 2018-08-28 Moa Lee , Joon Hyuk Chang

Text and vision foundation models can perform many tasks in a zero-shot setting, a desirable property that enables these systems to be applied in general and low-resource settings. There has been far less work, however, on the zero-shot…

Computation and Language · Computer Science 2024-03-29 Rao Ma , Adian Liusie , Mark J. F. Gales , Kate M. Knill

Dysarthric speech reconstruction (DSR) systems aim to automatically convert dysarthric speech into normal-sounding speech. The technology eases communication with speakers affected by the neuromotor disorder and enhances their social…

Sound · Computer Science 2024-01-29 Yuejiao Wang , Xixin Wu , Disong Wang , Lingwei Meng , Helen Meng

We developed dysarthric speech intelligibility classifiers on 551,176 disordered speech samples contributed by a diverse set of 468 speakers, with a range of self-reported speaking disorders and rated for their overall intelligibility on a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-17 Subhashini Venugopalan , Jimmy Tobin , Samuel J. Yang , Katie Seaver , Richard J. N. Cave , Pan-Pan Jiang , Neil Zeghidour , Rus Heywood , Jordan Green , Michael P. Brenner

Training machine learning algorithms for speech applications requires large, labeled training data sets. This is problematic for clinical applications where obtaining such data is prohibitively expensive because of privacy concerns or lack…

Audio and Speech Processing · Electrical Eng. & Systems 2018-04-30 Yishan Jiao , Ming Tu , Visar Berisha , Julie Liss