中文
相关论文

相关论文: Intel Labs at Ego4D Challenge 2022: A Better Basel…

200 篇论文

Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a speaker embedding…

音频与语音处理 · 电气工程与系统科学 2024-05-16 Jenthe Thienpondt , Kris Demuynck

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

计算机视觉与模式识别 · 计算机科学 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

Voice activity detection (VAD) is an essential pre-processing step for tasks such as automatic speech recognition (ASR) and speaker recognition. A basic goal is to remove silent segments within an audio, while a more general VAD system…

音频与语音处理 · 电气工程与系统科学 2020-09-22 Yefei Chen , Shuai Wang , Yanmin Qian , Kai Yu

Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the temporal correspondence of audiovisual cues, such as the…

多媒体 · 计算机科学 2025-02-11 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

Voice activity detection is the task of detecting speech regions in a given audio stream or recording. First, we design a neural network combining trainable filters and recurrent layers to tackle voice activity detection directly from the…

音频与语音处理 · 电气工程与系统科学 2020-05-27 Marvin Lavechin , Marie-Philippe Gill , Ruben Bousbib , Hervé Bredin , Leibny Paola Garcia-Perera

This paper describes the ByteDance speaker diarization system for the fourth track of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). The VoxSRC-21 provides both the dev set and test set of VoxConverse for use in validation and…

声音 · 计算机科学 2021-09-07 Keke Wang , Xudong Mao , Hao Wu , Chen Ding , Chuxiang Shang , Rui Xia , Yuxuan Wang

This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Unified Context Network…

计算机视觉与模式识别 · 计算机科学 2022-06-23 Yuanhang Zhang , Susan Liang , Shuang Yang , Shiguang Shan

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

In this paper we provide the technique report of Ego4D natural language query challenge in CVPR 2022. Natural language query task is challenging due to the requirement of comprehensive understanding of video contents. Most previous works…

计算机视觉与模式识别 · 计算机科学 2022-08-11 Sipeng Zheng , Qi Zhang , Bei Liu , Qin Jin , Jianlong Fu

This report presents the system developed by the ABSP Laboratory team for the third DIHARD speech diarization challenge. Our main contribution in this work is to develop a simple and efficient solution for acoustic domain dependent speech…

声音 · 计算机科学 2021-01-26 A Kishore Kumar , Shefali Waldekar , Goutam Saha , Md Sahidullah

This technical report describes our system for track 1, 2 and 4 of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). By combining several ResNet variants, our submission for track 1 attained a minDCF of 0:090 with EER 1:401%. By…

音频与语音处理 · 电气工程与系统科学 2022-09-26 Qutang Cai , Guoqiang Hong , Zhijian Ye , Ximin Li , Haizhou Li

Audiovisual active speaker detection (ASD) in egocentric recordings is challenged by frequent occlusions, motion blur, and audio interference, which undermine the discernability of temporal synchrony between lip movement and speech.…

多媒体 · 计算机科学 2025-08-15 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

Active speaker detection (ASD) in egocentric videos presents unique challenges due to unstable viewpoints, motion blur, and off-screen speech sources - conditions under which traditional visual-centric methods degrade significantly. We…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yu Wang , Juhyung Ha , David J. Crandall

Voice Activity Detection (VAD) in the presence of background noise remains a challenging problem in speech processing. Accurate VAD is essential in automatic speech recognition, voice-to-text, conversational agents, etc, where noise can…

音频与语音处理 · 电气工程与系统科学 2025-07-31 Hamed Jafarzadeh Asl , Mahsa Ghazvini Nejad , Amin Edraki , Masoud Asgharian , Vahid Partovi Nia

Voice Activity Detection (VAD) is an important pre-processing step in a wide variety of speech processing systems. VAD should in a practical application be able to detect speech in both noisy and noise-free environments, while not…

音频与语音处理 · 电气工程与系统科学 2022-07-06 Claus Meyer Larsen , Peter Koch , Zheng-Hua Tan

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition,…

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple candidate speakers are visible this traditionally requires…

音频与语音处理 · 电气工程与系统科学 2022-05-12 Otavio Braga , Olivier Siohan

In this paper, we present the speaker diarization system for the Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) from team DKU_DukeECE. As the highly overlapped speech exists in the dataset, we employ an x-vector-based…

音频与语音处理 · 电气工程与系统科学 2022-02-08 Weiqing Wang , Xiaoyi Qin , Ming Li

With the rise in multimedia content over the years, more variety is observed in the recording environments of audio. An audio processing system might benefit when it has a module to identify the acoustic domain at its front-end. In this…

声音 · 计算机科学 2022-08-09 A Kishore Kumar , Shefali Waldekar , Md Sahidullah , Goutam Saha

In this paper we demonstrate that performance of voice activity detection (VAD) system operating in presence of background noise can be improved by concatenating acoustic input features with electroencephalography (EEG) features. We also…

声音 · 计算机科学 2020-03-18 Gautam Krishna , Co Tran , Mason Carnahan , Yan Han , Ahmed H Tewfik