English
Related papers

Related papers: Uncovering Voice Misuse Using Symbolic Mismatch

200 papers

Project Euphonia, a Google initiative, is dedicated to improving automatic speech recognition (ASR) of disordered speech. A central objective of the project is to create a large, high-quality, and diverse speech corpus. This report…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Pan-Pan Jiang , Jimmy Tobin , Katrin Tomanek , Robert L. MacDonald , Katie Seaver , Richard Cave , Marilyn Ladewig , Rus Heywood , Jordan R. Green

This paper presents a macroscopic approach to automatic detection of speech sound disorder (SSD) in child speech. Typically, SSD is manifested by persistent articulation and phonological errors on specific phonemes in the language. The…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-30 Si-Ioi Ng , Cymie Wing-Yee Ng , Jiarui Wang , Tan Lee

In speech production research, different imaging modalities have been employed to obtain accurate information about the movement and shaping of the vocal tract. Ultrasound is an affordable and non-invasive imaging modality with relatively…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Tamás Gábor Csapó , Kele Xu

The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these challenges by showcasing the utilization of automatic speech…

Major depressive disorder is a common mental disorder that affects almost 7% of the adult U.S. population. The 2017 Audio/Visual Emotion Challenge (AVEC) asks participants to build a model to predict depression levels based on the audio,…

Computation and Language · Computer Science 2018-03-29 Yuan Gong , Christian Poellabauer

Estimating personal well-being draws increasing attention particularly from healthcare and pharmaceutical industries. We propose an approach to estimate personal well-being in terms of various measurements such as anxiety, sleep quality and…

Computation and Language · Computer Science 2019-10-23 Samuel Kim , Namhee Kwon , Henry O'Connell

Many companies, including Google, Amazon, and Apple, offer voice assistants as a convenient solution for answering general voice queries and accessing their services. These voice assistants have gained popularity and can be easily accessed…

Human-Computer Interaction · Computer Science 2024-09-16 Tina Khezresmaeilzadeh , Elaine Zhu , Kiersten Grieco , Daniel J. Dubois , Konstantinos Psounis , David Choffnes

We developed dysarthric speech intelligibility classifiers on 551,176 disordered speech samples contributed by a diverse set of 468 speakers, with a range of self-reported speaking disorders and rated for their overall intelligibility on a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-17 Subhashini Venugopalan , Jimmy Tobin , Samuel J. Yang , Katie Seaver , Richard J. N. Cave , Pan-Pan Jiang , Neil Zeghidour , Rus Heywood , Jordan Green , Michael P. Brenner

In this article a DNN-based system for detection of three common voice disorders (vocal nodules, polyps and cysts; laryngeal neoplasm; unilateral vocal paralysis) is presented. The input to the algorithm is (at least 3-second long) audio…

Dementia is a neurodegenerative disease that causes gradual cognitive impairment, which is very common in the world and undergoes a lot of research every year to prevent and cure it. It severely impacts the patient's ability to remember…

We present ChildVox, a novel benchmark for characterizing the diverse acoustic signals through which children communicate. Specifically, ChildVox follows the full developmental trajectory from birth through school age, covering…

Mild Cognitive Impairment (MCI) is an early stage of Alzheimer's disease (AD), a form of neurodegenerative disorder. Early identification of MCI is crucial for delaying its progression through timely interventions. Existing research has…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-18 Kristin Qi , Jiatong Shi , Caroline Summerour , John A. Batsis , Xiaohui Liang

Speech deepfakes are artificial voices generated by machine learning models. Previous literature has highlighted deepfakes as one of the biggest security threats arising from progress in artificial intelligence due to their potential for…

Human-Computer Interaction · Computer Science 2023-08-04 Kimberly T. Mai , Sergi D. Bray , Toby Davies , Lewis D. Griffin

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

Sound · Computer Science 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

Physiological signals can potentially be applied as objective measures to understand the behavior and engagement of users interacting with information access systems. However, the signals are highly sensitive, and many controls are required…

Information Retrieval · Computer Science 2023-04-27 Kaixin Ji , Damiano Spina , Danula Hettiachchi , Flora Dilys Salim , Falk Scholer

With the availability of voice-enabled devices such as smart phones, mental health disorders could be detected and treated earlier, particularly post-pandemic. The current methods involve extracting features directly from audio signals. In…

Machine Learning · Computer Science 2022-05-17 Nasser Ghadiri , Rasoul Samani , Fahime Shahrokh

Stress is a major threat to well-being that manifests in a variety of physiological and mental symptoms. Utilising speech samples collected while the subject is undergoing an induced stress episode has recently shown promising results for…

Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Shota Horiguchi , Takanori Ashihara , Marc Delcroix , Atsushi Ando , Naohiro Tawara

We introduce a sophisticated multi-speaker speech data simulator, specifically engineered to generate multi-speaker speech recordings. A notable feature of this simulator is its capacity to modulate the distribution of silence and overlap…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-20 Tae Jin Park , He Huang , Coleman Hooper , Nithin Koluguri , Kunal Dhawan , Ante Jukic , Jagadeesh Balam , Boris Ginsburg

Speech-based depression detection has shown promise as an objective diagnostic tool, yet the cross-linguistic robustness of acoustic markers and their neurobiological underpinnings remain underexplored. This study extends Cross-Data…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-07 Fuxiang Tao , Dongwei Li , Shuning Tang , Xuri Ge , Wei Ma , Anna Esposito , Alessandro Vinciarelli
‹ Prev 1 4 5 6 7 8 10 Next ›