English
Related papers

Related papers: Personal VAD 2.0: Optimizing Personal Voice Activi…

200 papers

Speech Activity Detection (SAD) systems often misclassify singing as speech, leading to degraded performance in applications such as dialogue enhancement and automatic speech recognition. We introduce Singing-Robust Speech Activity…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-11 Philipp Grundhuber , Mhd Modar Halimeh , Martin Strauß , Emanuël A. P. Habets

The recent ubiquitous adoption of remote conferencing has been accompanied by omnipresent frustration with distorted or otherwise unclear voice communication. Audio enhancement can compensate for low-quality input signals from, for example,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-06 Philipp Schilk , Niccolò Polvani , Andrea Ronco , Milos Cernak , Michele Magno

This paper presents a simulation-based approach to own voice detection (OVD) in hearing aids using a single microphone. While OVD can significantly improve user comfort and speech intelligibility, existing solutions often rely on multiple…

Many existing works on voice conversion (VC) tasks use automatic speech recognition (ASR) models for ensuring linguistic consistency between source and converted samples. However, for the low-data resource domains, training a high-quality…

Sound · Computer Science 2023-05-25 Mayank Kumar Singh , Naoya Takahashi , Onoe Naoyuki

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

Multimedia · Computer Science 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

Recent advancements in Self-Supervised Learning (SSL) have shown promising results in Speaker Verification (SV). However, narrowing the performance gap with supervised systems remains an ongoing challenge. Several studies have observed that…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-25 Victor Miara , Theo Lepage , Reda Dehak

Personalizing Automatic Speech Recognition (ASR) for non-normative speech remains challenging because data collection is labor-intensive and model training is technically complex. To address these limitations, we propose Adapt4Me, a…

Human-Computer Interaction · Computer Science 2026-03-23 Niclas Pokel , Yiming Zhao , Pehuén Moure , Yingqiang Gao , Roman Böhringer

Deploying high-quality automatic speech recognition (ASR) on edge devices requires models that jointly optimize accuracy, latency, and memory footprint while operating entirely on CPU without GPU acceleration. We conduct a systematic…

Artificial Intelligence · Computer Science 2026-04-21 Nenad Banfic , David Fan , Kunal Vaishnavi , Sam Kemp , Sunghoon Choi , Rui Ren , Sayan Shaw , Meng Tang

Self-supervised learning (SSL) is a powerful tool that allows learning of underlying representations from unlabeled data. Transformer based models such as wav2vec 2.0 and HuBERT are leading the field in the speech domain. Generally these…

Computation and Language · Computer Science 2022-02-08 Bethan Thomas , Samuel Kessler , Salah Karout

Voice-based human-machine interfaces with an automatic speaker verification (ASV) component are commonly used in the market. However, the threat from presentation attacks is also growing since attackers can use recent speech synthesis…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-11 Xin Wang , Junichi Yamagishi

Voice activity detection is the task of detecting speech regions in a given audio stream or recording. First, we design a neural network combining trainable filters and recurrent layers to tackle voice activity detection directly from the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-27 Marvin Lavechin , Marie-Philippe Gill , Ruben Bousbib , Hervé Bredin , Leibny Paola Garcia-Perera

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Otavio Braga , Takaki Makino , Olivier Siohan , Hank Liao

One in four people dementia live alone, leading family members to take on caregiving roles from a distance. Many researchers have developed remote monitoring solutions to lessen caregiving needs; however, limitations remain including…

Sound · Computer Science 2025-09-01 Dong Yoon Lee , Alyssa Weakley , Hui Wei , Blake Brown , Keyana Carrion , Shijia Pan

Visual speech recognition (VSR) is the task of recognizing spoken language from video input only, without any audio. VSR has many applications as an assistive technology, especially if it could be deployed in mobile devices and embedded…

Computation and Language · Computer Science 2019-06-06 Nilay Shrivastava , Astitwa Saxena , Yaman Kumar , Rajiv Ratn Shah , Debanjan Mahata , Amanda Stent

Reducing the latency and model size has always been a significant research problem for live Automatic Speech Recognition (ASR) application scenarios. Along this direction, model quantization has become an increasingly popular approach to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-06 Shaojin Ding , Phoenix Meadowlark , Yanzhang He , Lukasz Lew , Shivani Agrawal , Oleg Rybakov

Audio-based automatic speech recognition (ASR) degrades significantly in noisy environments and is particularly vulnerable to interfering speech, as the model cannot determine which speaker to transcribe. Audio-visual speech recognition…

Sound · Computer Science 2022-07-18 Bowen Shi , Wei-Ning Hsu , Abdelrahman Mohamed

Speaker Recognition and Speaker Identification are challenging tasks with essential applications such as automation, authentication, and security. Deep learning approaches like SincNet and AM-SincNet presented great results on these tasks.…

Sound · Computer Science 2020-10-20 João Antônio Chagas Nunes , David Macêdo , Cleber Zanchettin

Voice activity detection (VAD) is an important pre-processing step for speech technology applications. The task consists of deriving segment boundaries of audio signals which contain voicing information. In recent years, it has been shown…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-28 Eklavya Sarkar , RaviShankar Prasad , Mathew Magimai. -Doss

Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven…

Sound · Computer Science 2025-08-22 Dena Mujtaba , Nihar Mahapatra

Self-supervised learned models have been found to be very effective for certain speech tasks such as automatic speech recognition, speaker identification, keyword spotting and others. While the features are undeniably useful in speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-05 Ravi Shankar , Ke Tan , Buye Xu , Anurag Kumar