English
Related papers

Related papers: Optimal Transport-based Adaptation in Dysarthric S…

200 papers

Dysarthria, a common issue among stroke patients, severely impacts speech intelligibility. Inappropriate pauses are crucial indicators in severity assessment and speech-language therapy. We propose to extend a large-scale speech recognition…

Computation and Language · Computer Science 2024-03-01 Jeehyun Lee , Yerin Choi , Tae-Jin Song , Myoung-Wan Koo

Disordered speech recognition is a highly challenging task. The underlying neuro-motor conditions of people with speech disorders, often compounded with co-occurring physical disabilities, lead to the difficulty in collecting large…

Sound · Computer Science 2022-01-20 Mengzhe Geng , Xurong Xie , Shansong Liu , Jianwei Yu , Shoukang Hu , Xunying Liu , Helen Meng

Although deep learning-based end-to-end Automatic Speech Recognition (ASR) has shown remarkable performance in recent years, it suffers severe performance regression on test samples drawn from different data distributions. Test-time…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-22 Guan-Ting Lin , Shang-Wen Li , Hung-yi Lee

Speech recognition systems have improved dramatically over the last few years, however, their performance is significantly degraded for the cases of accented or impaired speech. This work explores domain adversarial neural networks (DANN)…

Sound · Computer Science 2020-10-09 Dominika Woszczyk , Stavros Petridis , David Millard

Domain adaptation arises as an important problem in statistical learning theory when the data-generating processes differ between training and test samples, respectively called source and target domains. Recent theoretical advances show…

Machine Learning · Statistics 2022-10-25 Mourad El Hamri , Younès Bennani , Issam Falih

Dysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-26 Dimme de Groot , Tanvina Patel , Devendra Kayande , Odette Scharenborg , Zhengjun Yue

Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance arising from long-tailed direction-of-arrival (DoA)…

Sound · Computer Science 2026-01-27 Zexia Fan , Yu Chen , Qiquan Zhang , Kainan Chen , Xinyuan Qian

Data selection techniques applied to neural machine translation (NMT) aim to increase the performance of a model by retrieving a subset of sentences for use as training data. One of the possible data selection techniques are transductive…

Computation and Language · Computer Science 2018-11-08 Alberto Poncelas , Gideon Maillette de Buy Wenniger , Andy Way

We propose a separation guided speaker diarization (SGSD) approach by fully utilizing a complementarity of speech separation and speaker clustering. Since the conventional clustering-based speaker diarization (CSD) approach cannot well…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-07 Shu-Tong Niu , Jun Du , Lei Sun , Chin-Hui Lee

Domain adaptation (DA) is transfer learning which aims to leverage labeled data in a related source domain to achieve informed knowledge transfer and help the classification of unlabeled data in a target domain. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2017-05-25 Lingkun Luo , Xiaofang Wang , Shiqiang Hu , Liming Chen

Unsupervised spoken term discovery (UTD) aims at finding recurring segments of speech from a corpus of acoustic speech data. One potential approach to this problem is to use dynamic time warping (DTW) to find well-aligning patterns from the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Okko Räsänen , María Andrea Cruz Blandón

Evaluating natural language generation models, particularly for method name prediction, poses significant challenges. A robust metric must account for the versatility of method naming, considering both semantic and syntactic variations.…

Computation and Language · Computer Science 2024-08-14 Ravil Mussabayev

Distributional shifts between training and inference time data remain a central challenge in machine learning, often leading to poor performance. It motivated the study of principled approaches for domain alignment, such as optimal…

Machine Learning · Computer Science 2026-03-09 Abdel Djalil Sad Saoud , Fred Maurice Ngolè Mboula , Hanane Slimani

Systematics contaminate observables, leading to distribution shifts relative to theoretically simulated signals-posing a major challenge for using pre-trained models to label such observables. Since systematics are often poorly understood…

Instrumentation and Methods for Astrophysics · Physics 2025-11-18 Sultan Hassan , Sambatra Andrianomena , Benjamin D. Wandelt

Personalizing dysarthric ASR is hindered by demanding enrollment collection and per-user training. We propose a hybrid meta-training method for a single model, enabling zero-shot and few-shot on-the-fly personalization via in-context…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-24 Dhruuv Agarwal , Harry Zhang , Yang Yu , Quan Wang

In this paper, we propose a domain adversarial training (DAT) algorithm to alleviate the accented speech recognition problem. In order to reduce the mismatch between labeled source domain data ("standard" accent) and unlabeled target domain…

Computation and Language · Computer Science 2018-06-08 Sining Sun , Ching-Feng Yeh , Mei-Yuh Hwang , Mari Ostendorf , Lei Xie

Machine learning algorithms typically assume that the training and test samples come from the same distributions, i.e., in-distribution. However, in open-world scenarios, streaming big data can be Out-Of-Distribution (OOD), rendering these…

Machine Learning · Computer Science 2022-11-10 Anique Tahir , Lu Cheng , Ruocheng Guo , Huan Liu

Dysarthric speech recognition has posed major challenges due to lack of training data and heavy mismatch in speaker characteristics. Recent ASR systems have benefited from readily available pretrained models such as wav2vec2 to improve the…

The disparity in phonology between learner's native (L1) and target (L2) language poses a significant challenge for mispronunciation detection and diagnosis (MDD) systems. This challenge is further intensified by lack of annotated L2 data.…

Sound · Computer Science 2023-08-08 Yassine El Kheir , Shammur Absar Chowdhury , Ahmed Ali

Optimal transport distances have found many applications in machine learning for their capacity to compare non-parametric probability distributions. Yet their algorithmic complexity generally prevents their direct use on large scale…

Machine Learning · Computer Science 2021-03-08 Kilian Fatras , Thibault Séjourné , Nicolas Courty , Rémi Flamary