English
Related papers

Related papers: Semantic VAD: Low-Latency Voice Activity Detection…

200 papers

In this work, we propose novel decoding algorithms to enable streaming automatic speech recognition (ASR) on unsegmented long-form recordings without voice activity detection (VAD), based on monotonic chunkwise attention (MoChA) with an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-16 Hirofumi Inaguma , Tatsuya Kawahara

In this paper, we show how to use audio to supervise the learning of active speaker detection in video. Voice Activity Detection (VAD) guides the learning of the vision-based classifier in a weakly supervised manner. The classifier uses…

Computer Vision and Pattern Recognition · Computer Science 2016-03-30 Punarjay Chakravarty , Tinne Tuytelaars

Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainly due to the limited ability to handle overlapping speech.…

Speech activity detection (SAD) is an essential component for a variety of speech processing applications. It has been observed that performances of various speech based tasks are very much dependent on the efficiency of the SAD. In this…

Multimedia · Computer Science 2012-10-09 Md. Sahidullah , Goutam Saha

Past studies on end-to-end meeting transcription have focused on model architecture and have mostly been evaluated on simulated meeting data. We present a novel study aiming to optimize the use of a Speaker-Attributed ASR (SA-ASR) system in…

Computation and Language · Computer Science 2024-09-06 Can Cui , Imran Ahamad Sheikh , Mostafa Sadeghi , Emmanuel Vincent

We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phrase. We first demonstrate that by including more audio…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-03 Siddharth Sigtia , John Bridle , Hywel Richards , Pascal Clark , Erik Marchi , Vineet Garg

Voice activity and overlapped speech detection (respectively VAD and OSD) are key pre-processing tasks for speaker diarization. The final segmentation performance highly relies on the robustness of these sub-tasks. Recent studies have shown…

Recent audio generation models typically rely on Variational Autoencoders (VAEs) and perform generation within the VAE latent space. Although VAEs excel at compression and reconstruction, their latents inherently encode low-level acoustic…

Sound · Computer Science 2026-02-27 Zeyu Xie , Chenxing Li , Qiao Jin , Xuenan Xu , Guanrou Yang , Wenfu Wang , Mengyue Wu , Dong Yu , Yuexian Zou

We present a novel personalized voice activity detection (PVAD) learning method that does not require enrollment data during training. PVAD is a task to detect the speech segments of a specific target speaker at the frame level using…

Audio-visual automatic speech recognition (AV-ASR) introduces the video modality into the speech recognition process, often by relying on information conveyed by the motion of the speaker's mouth. The use of the video signal requires…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Dmitriy Serdyuk , Otavio Braga , Olivier Siohan

This work proposes a novel support vector machine (SVM) based robust automatic speech recognition (ASR) front-end that operates on an ensemble of the subband components of high-dimensional acoustic waveforms. The key issues of selecting the…

Computation and Language · Computer Science 2014-01-15 Jibran Yousafzai , Zoran Cvetkovic , Peter Sollich , Matthew Ager

We describe here our work with automatic speech recognition (ASR) in the context of voice search functionality on the Flipkart e-Commerce platform. Starting with the deep learning architecture of Listen-Attend-Spell (LAS), we build upon and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-01 Raviraj Joshi , Venkateshan Kannan

Voice Activity Detection (VAD) is a fundamental module in many audio applications. Recent state-of-the-art VAD systems are often based on neural networks, but they require a computational budget that usually exceeds the capabilities of a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-07 Niccolo' Polvani , Damien Ronssin , Milos Cernak

Video Anomaly Detection (VAD) is a fundamental challenge in computer vision, particularly due to the open-set nature of anomalies. While recent training-free approaches utilizing Vision-Language Models (VLMs) have shown promise, they…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Lokman Bekit , Hamza Karim , Nghia T Nguyen , Yasin Yilmaz

Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting the presence of speech from a known target-speaker in an audio frame. Recently, deep neural network-based models have shown good performance in this task. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-07 Holger Severin Bovbjerg , Jan Østergaard , Jesper Jensen , Zheng-Hua Tan

End-to-end automatic speech recognition systems represent the state of the art, but they rely on thousands of hours of manually annotated speech for training, as well as heavyweight computation for inference. Of course, this impedes…

Computation and Language · Computer Science 2022-11-22 Raphael Tang , Karun Kumar , Gefei Yang , Akshat Pandey , Yajie Mao , Vladislav Belyaev , Madhuri Emmadi , Craig Murray , Ferhan Ture , Jimmy Lin

Automatic speech recognition (ASR) is widely used in consumer electronics. ASR greatly improves the utility and accessibility of technology, but usually the output is only word sequences without punctuation. This can result in ambiguity in…

Computation and Language · Computer Science 2021-02-23 Andrew Silva , Barry-John Theobald , Nicholas Apostoloff

Active Speaker Detection (ASD) aims to identify who is speaking in each frame of a video. ASD reasons from audio and visual information from two contexts: long-term intra-speaker context and short-term inter-speaker context. Long-term…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Xizi Wang , Feng Cheng , Gedas Bertasius , David Crandall

Temporal action detection (TAD) aims to detect the semantic labels and boundaries of action instances in untrimmed videos. Current mainstream approaches are multi-step solutions, which fall short in efficiency and flexibility. In this…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Shimin Chen , Chen Chen , Wei Li , Xunqiang Tao , Yandong Guo

Voice controlled virtual assistants (VAs) are now available in smartphones, cars, and standalone devices in homes. In most cases, the user needs to first "wake-up" the VA by saying a particular word/phrase every time he or she wants the VA…

Human-Computer Interaction · Computer Science 2019-02-05 Atta Norouzian , Bogdan Mazoure , Dermot Connolly , Daniel Willett
‹ Prev 1 4 5 6 7 8 10 Next ›