English
Related papers

Related papers: Intel Labs at Ego4D Challenge 2022: A Better Basel…

200 papers

Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a speaker embedding…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-16 Jenthe Thienpondt , Kris Demuynck

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

Computer Vision and Pattern Recognition · Computer Science 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

Voice activity detection (VAD) is an essential pre-processing step for tasks such as automatic speech recognition (ASR) and speaker recognition. A basic goal is to remove silent segments within an audio, while a more general VAD system…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-22 Yefei Chen , Shuai Wang , Yanmin Qian , Kai Yu

Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the temporal correspondence of audiovisual cues, such as the…

Multimedia · Computer Science 2025-02-11 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

Voice activity detection is the task of detecting speech regions in a given audio stream or recording. First, we design a neural network combining trainable filters and recurrent layers to tackle voice activity detection directly from the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-27 Marvin Lavechin , Marie-Philippe Gill , Ruben Bousbib , Hervé Bredin , Leibny Paola Garcia-Perera

This paper describes the ByteDance speaker diarization system for the fourth track of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). The VoxSRC-21 provides both the dev set and test set of VoxConverse for use in validation and…

Sound · Computer Science 2021-09-07 Keke Wang , Xudong Mao , Hao Wu , Chen Ding , Chuxiang Shang , Rui Xia , Yuxuan Wang

This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Unified Context Network…

Computer Vision and Pattern Recognition · Computer Science 2022-06-23 Yuanhang Zhang , Susan Liang , Shuang Yang , Shiguang Shan

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

In this paper we provide the technique report of Ego4D natural language query challenge in CVPR 2022. Natural language query task is challenging due to the requirement of comprehensive understanding of video contents. Most previous works…

Computer Vision and Pattern Recognition · Computer Science 2022-08-11 Sipeng Zheng , Qi Zhang , Bei Liu , Qin Jin , Jianlong Fu

This report presents the system developed by the ABSP Laboratory team for the third DIHARD speech diarization challenge. Our main contribution in this work is to develop a simple and efficient solution for acoustic domain dependent speech…

Sound · Computer Science 2021-01-26 A Kishore Kumar , Shefali Waldekar , Goutam Saha , Md Sahidullah

This technical report describes our system for track 1, 2 and 4 of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). By combining several ResNet variants, our submission for track 1 attained a minDCF of 0:090 with EER 1:401%. By…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-26 Qutang Cai , Guoqiang Hong , Zhijian Ye , Ximin Li , Haizhou Li

Audiovisual active speaker detection (ASD) in egocentric recordings is challenged by frequent occlusions, motion blur, and audio interference, which undermine the discernability of temporal synchrony between lip movement and speech.…

Multimedia · Computer Science 2025-08-15 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

Active speaker detection (ASD) in egocentric videos presents unique challenges due to unstable viewpoints, motion blur, and off-screen speech sources - conditions under which traditional visual-centric methods degrade significantly. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yu Wang , Juhyung Ha , David J. Crandall

Voice Activity Detection (VAD) in the presence of background noise remains a challenging problem in speech processing. Accurate VAD is essential in automatic speech recognition, voice-to-text, conversational agents, etc, where noise can…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-31 Hamed Jafarzadeh Asl , Mahsa Ghazvini Nejad , Amin Edraki , Masoud Asgharian , Vahid Partovi Nia

Voice Activity Detection (VAD) is an important pre-processing step in a wide variety of speech processing systems. VAD should in a practical application be able to detect speech in both noisy and noise-free environments, while not…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-06 Claus Meyer Larsen , Peter Koch , Zheng-Hua Tan

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition,…

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple candidate speakers are visible this traditionally requires…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Otavio Braga , Olivier Siohan

In this paper, we present the speaker diarization system for the Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) from team DKU_DukeECE. As the highly overlapped speech exists in the dataset, we employ an x-vector-based…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-08 Weiqing Wang , Xiaoyi Qin , Ming Li

With the rise in multimedia content over the years, more variety is observed in the recording environments of audio. An audio processing system might benefit when it has a module to identify the acoustic domain at its front-end. In this…

Sound · Computer Science 2022-08-09 A Kishore Kumar , Shefali Waldekar , Md Sahidullah , Goutam Saha

In this paper we demonstrate that performance of voice activity detection (VAD) system operating in presence of background noise can be improved by concatenating acoustic input features with electroencephalography (EEG) features. We also…

Sound · Computer Science 2020-03-18 Gautam Krishna , Co Tran , Mason Carnahan , Yan Han , Ahmed H Tewfik