English
Related papers

Related papers: Automated speech audiometry: Can it work using ope…

200 papers

Early detection and treatment of depression is essential in promoting remission, preventing relapse, and reducing the emotional burden of the disease. Current diagnoses are primarily subjective, inconsistent across professionals, and…

Machine Learning · Computer Science 2020-02-03 Karol Chlasta , Krzysztof Wołk , Izabela Krejtz

Always-on spoken language interfaces, e.g. personal digital assistants, rely on a wake word to start processing spoken input. We present novel methods to train a hybrid DNN/HMM wake word detection system from partially labeled training…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-30 Yiming Wang , Hang Lv , Daniel Povey , Lei Xie , Sanjeev Khudanpur

Personalizing automatic speech recognition (ASR) systems for non-normative speech, such as dysarthric and aphasic speech, is challenging. While speaker-specific fine-tuning (SS-FT) is widely used, it is typically initialized directly from a…

Sound · Computer Science 2026-03-17 Shan Jiang , Jiawen Qi , Chuanbing Huo , Yingqiang Gao , Qinyu Chen

The popularity of automatic speech recognition (ASR) systems nowadays leads to an increasing need for improving their accessibility. Handling stuttering speech is an important feature for accessible ASR systems. To improve the accessibility…

Sound · Computer Science 2023-08-31 Yi Liu , Yuekang Li , Gelei Deng , Felix Juefei-Xu , Yao Du , Cen Zhang , Chengwei Liu , Yeting Li , Lei Ma , Yang Liu

Automatic coded audio quality predictors are typically designed for evaluating single channels without considering any spatial aspects. With InSE-NET [1], we demonstrated mimicking a state-of-the-art coded audio quality metric (ViSQOL-v3…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-26 Arijit Biswas , Guanxin Jiang

Due to the mismatch of statistical distributions of acoustic speech between training and testing sets, the performance of spoken language identification (SLID) could be drastically degraded. In this paper, we propose an unsupervised neural…

Machine Learning · Computer Science 2020-12-25 Xugang Lu , Peng Shen , Yu Tsao , Hisashi Kawai

We describe the neural-network training framework used in the Kaldi speech recognition toolkit, which is geared towards training DNNs with large amounts of training data using multiple GPU-equipped or multi-core machines. In order to be as…

Neural and Evolutionary Computing · Computer Science 2015-06-24 Daniel Povey , Xiaohui Zhang , Sanjeev Khudanpur

We address zero-shot TTS systems' noise-robustness problem by proposing a dual-objective training for the speaker encoder using self-supervised DINO loss. This approach enhances the speaker encoder with the speech synthesis objective,…

In this study, listeners of varied Indian nativities are asked to listen and recognize TIMIT utterances spoken by American speakers. We have three kinds of responses from each listener while they recognize an utterance: 1. Sentence…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-09 Abhayjeet Singh , Achuth Rao MV , Rakesh Vaideeswaran , Chiranjeevi Yarra , Prasanta Kumar Ghosh

In recent years, speech-based self-supervised learning (SSL) has made significant progress in various tasks, including automatic speech recognition (ASR). An ASR model with decent performance can be realized by fine-tuning an SSL model with…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-30 Zhisheng Zheng , Ziyang Ma , Yu Wang , Xie Chen

Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information. This comprehensive nature of speech significantly impacts communication and is crucial for human-computer…

Computation and Language · Computer Science 2025-01-17 Junyi Ao , Yuancheng Wang , Xiaohai Tian , Dekun Chen , Jun Zhang , Lu Lu , Yuxuan Wang , Haizhou Li , Zhizheng Wu

The goal of this paper is to simulate the benefits of jointly applying active learning (AL) and semi-supervised training (SST) in a new speech recognition application. Our data selection approach relies on confidence filtering, and its…

Computation and Language · Computer Science 2019-03-08 Thomas Drugman , Janne Pylkkonen , Reinhard Kneser

Auscultation is crucial for diagnosing lung diseases. The COVID-19 pandemic has revealed the limitations of traditional, in-person lung sound assessments. To overcome these issues, advancements in digital stethoscopes and artificial…

Sound · Computer Science 2025-05-30 Seung Gyu Jeong , Seong Eun Kim

State-of-the-art speaker recognition systems are trained with a large amount of human-labeled training data set. Such a training set is usually composed of various data sources to enhance the modeling capability of models. However, in…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-03 Rongjin Li , Weibin Zhang , Dongpeng Chen

One of the established multi-lingual methods for testing speech intelligibility is the matrix sentence test (MST). Most versions of this test are designed with audio-only stimuli. Nevertheless, visual cues play an important role in speech…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-01 Gerard Llorach , Frederike Kirschner , Giso Grimm , Melanie A. Zokoll , Kirsten C. Wagener , Volker Hohmann

Automatic Speech Recognition (ASR) systems are commonly evaluated using aggregate metrics such as Word Error Rate (WER), which do not capture the linguistic structure of errors. Fine-grained analysis, such as Part-of-Speech (PoS)-wise error…

Computation and Language · Computer Science 2026-05-28 Prasenjit K Mudi , Dahlia Devapriya , Sheetal Kalyani

Spoken language recognition (SLR) refers to the automatic process used to determine the language present in a speech sample. SLR is an important task in its own right, for example, as a tool to analyze or categorize large amounts of…

Computation and Language · Computer Science 2022-08-15 Luciana Ferrer , Diego Castan , Mitchell McLaren , Aaron Lawson

This work describes the speaker verification system developed by Human Language Technology Laboratory, National University of Singapore (HLT-NUS) for 2019 NIST Multimedia Speaker Recognition Evaluation (SRE). The multimedia research has…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-09 Rohan Kumar Das , Ruijie Tao , Jichen Yang , Wei Rao , Cheng Yu , Haizhou Li

There is increasingly more evidence that automatic speech recognition (ASR) systems are biased against different speakers and speaker groups, e.g., due to gender, age, or accent. Research on bias in ASR has so far primarily focused on…

Computation and Language · Computer Science 2025-07-09 Tanvina Patel , Wiebke Hutiri , Aaron Yi Ding , Odette Scharenborg

Self-supervised learning (SSL) methods have proven to be very successful in automatic speech recognition (ASR). These great improvements have been reported mostly based on highly curated datasets such as LibriSpeech for non-streaming…

Sound · Computer Science 2022-05-19 Mostafa Karimi , Changliang Liu , Kenichi Kumatani , Yao Qian , Tianyu Wu , Jian Wu