English
Related papers

Related papers: End-to-end Audiovisual Speech Activity Detection w…

200 papers

In automatic speech recognition (ASR), wideband (WB) and narrowband (NB) speech signals with different sampling rates typically use separate acoustic models. Therefore mixed-bandwidth (MB) acoustic modeling has important practical values…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-12 Khoi-Nguyen C. Mac , Xiaodong Cui , Wei Zhang , Michael Picheny

Reverberation is present in our workplaces, our homes, concert halls and theatres. This paper investigates how deep learning can use the effect of reverberation on speech to classify a recording in terms of the room in which it was…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-03 Constantinos Papayiannis , Christine Evers , Patrick A. Naylor

Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA)…

Sound · Computer Science 2024-03-22 Samuel Pegg , Kai Li , Xiaolin Hu

Voice Activity Detection (VAD) is a fundamental module in many audio applications. Recent state-of-the-art VAD systems are often based on neural networks, but they require a computational budget that usually exceeds the capabilities of a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-07 Niccolo' Polvani , Damien Ronssin , Milos Cernak

Sociometric badges are an emerging technology for study how teams interact in physical places. Audio data recorded by sociometric badges is often downsampled to not record discussions of the sociometric badges holders. To gain more…

End-to-end neural network systems for automatic speech recognition (ASR) are trained from acoustic features to text transcriptions. In contrast to modular ASR systems, which contain separately-trained components for acoustic modeling,…

Computation and Language · Computer Science 2020-04-21 Yonatan Belinkov , Ahmed Ali , James Glass

In this paper we demonstrate that performance of voice activity detection (VAD) system operating in presence of background noise can be improved by concatenating acoustic input features with electroencephalography (EEG) features. We also…

Sound · Computer Science 2020-03-18 Gautam Krishna , Co Tran , Mason Carnahan , Yan Han , Ahmed H Tewfik

In this work, we propose a novel cross-talk rejection framework for a multi-channel multi-talker setup for a live multiparty interactive show. Our far-field audio setup is required to be hands-free during live interaction and comprises four…

Sound · Computer Science 2024-02-16 Hyewon Han , Naveen Kumar

Audio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, there is a noticeable…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-02 Ruijie Tao , Xinyuan Qian , Rohan Kumar Das , Xiaoxue Gao , Jiadong Wang , Haizhou Li

Voice activity detection (VAD) is a challenging task in low signal-to-noise ratio (SNR) environment, especially in non-stationary noise. To deal with this issue, we propose a novel attention module that can be integrated in Long Short-Term…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-26 Joohyung Lee , Youngmoon Jung , Hoirin Kim

It has been shown that the intelligibility of noisy speech can be improved by speech enhancement (SE) algorithms. However, monaural SE has not been established as an effective frontend for automatic speech recognition (ASR) in noisy…

Sound · Computer Science 2024-03-12 Yufeng Yang , Ashutosh Pandey , DeLiang Wang

Various neural network-based approaches have been proposed for more robust and accurate voice activity detection (VAD). Manual design of such neural architectures is an error-prone and time-consuming process, which prompted the development…

Sound · Computer Science 2022-10-05 Daniel Rho , Jinhyeok Park , Jong Hwan Ko

Recent achievements in end-to-end deep learning have encouraged the exploration of tasks dealing with highly structured data with unified deep network models. Having such models for compressing audio signals has been challenging since it…

Machine Learning · Computer Science 2021-07-14 Daniela N. Rim , Inseon Jang , Heeyoul Choi

Classroom activity detection (CAD) focuses on accurately classifying whether the teacher or student is speaking and recording both the length of individual utterances during a class. A CAD solution helps teachers get instant feedback on…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-12 Hang Li , Yu Kang , Wenbiao Ding , Song Yang , Songfan Yang , Gale Yan Huang , Zitao Liu

Predicting words and subword units (WSUs) as the output has shown to be effective for the attention-based encoder-decoder (AED) model in end-to-end speech recognition. However, as one input to the decoder recurrent neural network (RNN),…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-08 Zhong Meng , Yashesh Gaur , Jinyu Li , Yifan Gong

Estimating noise information exactly is crucial for noise aware training in speech applications including speech enhancement (SE) which is our focus in this paper. To estimate noise-only frames, we employ voice activity detection (VAD) to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-04 Joohyung Lee , Youngmoon Jung , Myunghun Jung , Hoirin Kim

A deep learning approach has been widely applied in sequence modeling problems. In terms of automatic speech recognition (ASR), its performance has significantly been improved by increasing large speech corpus and deeper neural network.…

Computation and Language · Computer Science 2016-12-28 Zewang Zhang , Zheng Sun , Jiaqi Liu , Jingwen Chen , Zhao Huo , Xiao Zhang

Voice activity detection (VAD) is essential for speech-driven applications, but remains far from perfect in noisy and resource-limited environments. Existing methods often lack robustness to noise, and their frame-wise classification losses…

Sound · Computer Science 2025-08-29 Chien-Chun Wang , En-Lun Yu , Jeih-Weih Hung , Shih-Chieh Huang , Berlin Chen

Conventional audio-visual approaches for active speaker detection (ASD) typically rely on visually pre-extracted face tracks and the corresponding single-channel audio to find the speaker in a video. Therefore, they tend to fail every time…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Davide Berghi , Philip J. B. Jackson

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

Sound · Computer Science 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie