English
Related papers

Related papers: AVA-ActiveSpeaker: An Audio-Visual Dataset for Act…

200 papers

Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Lionel Pibre , Francisco Madrigal , Cyrille Equoy , Frédéric Lerasle , Thomas Pellegrini , Julien Pinquier , Isabelle Ferrané

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

Visual voice activity detection (V-VAD) uses visual features to predict whether a person is speaking or not. V-VAD is useful whenever audio VAD (A-VAD) is inefficient either because the acoustic signal is difficult to analyze or because it…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Sylvain Guy , Stéphane Lathuilière , Pablo Mesejo , Radu Horaud

Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide complementary information for localization and tracking. With…

This paper delves into the challenging task of Active Speaker Detection (ASD), where the system needs to determine in real-time whether a person is speaking or not in a series of video frames. While previous works have made significant…

Computer Vision and Pattern Recognition · Computer Science 2024-09-16 Arnav Kundu , Yanzi Jin , Mohammad Sekhavat , Max Horton , Danny Tormoen , Devang Naik

In this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild. This dataset represents a significant expansion over the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-17 Yuke Lin , Ming Cheng , Fulin Zhang , Yingying Gao , Shilei Zhang , Ming Li

Robots are becoming everyday devices, increasing their interaction with humans. To make human-machine interaction more natural, cognitive features like Visual Voice Activity Detection (VVAD), which can detect whether a person is speaking or…

Computer Vision and Pattern Recognition · Computer Science 2024-01-03 Adrian Lubitz , Matias Valdenegro-Toro , Frank Kirchner

In recent years, significant progress has been made in automatic lip reading. But these methods require large-scale datasets that do not exist for many low-resource languages. In this paper, we have presented a new multipurpose audio-visual…

Computation and Language · Computer Science 2023-01-25 Javad Peymanfard , Samin Heydarian , Ali Lashini , Hossein Zeinali , Mohammad Reza Mohammadi , Nasser Mozayani

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 Rahul Sharma , Krishna Somandepalli , Shrikanth Narayanan

Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras and microphones are…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-05 Davide Berghi , Philip J. B. Jackson

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…

Active authentication refers to a new mode of identity verification in which biometric indicators are continuously tested to provide real-time or near real-time monitoring of an authorized access to a service or use of a device. This is in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-28 Zhong Meng , M Umair Bin Altaf , Biing-Hwang , Juang

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

Sound · Computer Science 2022-06-22 Yuan Gong , Jin Yu , James Glass

We propose a dataset, AVASpeech-SMAD, to assist speech and music activity detection research. With frame-level music labels, the proposed dataset extends the existing AVASpeech dataset, which originally consists of 45 hours of audio and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-03 Yun-Ning Hung , Karn N. Watcharasupat , Chih-Wei Wu , Iroro Orife , Kelian Li , Pavan Seshadri , Junyoung Lee

Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a…

Sound · Computer Science 2021-05-11 Heinrich Dinkel , Shuai Wang , Xuenan Xu , Mengyue Wu , Kai Yu

In this paper, we introduce a large-scale and high-quality audio-visual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-14 Yuke Lin , Xiaoyi Qin , Guoqing Zhao , Ming Cheng , Ning Jiang , Haiyang Wu , Ming Li

Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not explored and are not applicable for real-time applications…

Sound · Computer Science 2022-11-24 Fiseha B. Tesema , Zheyuan Lin , Shiqiang Zhu , Wei Song , Jason Gu , Hong Wu

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Otavio Braga , Takaki Makino , Olivier Siohan , Hank Liao

We address the problem of active speaker detection through a new framework, called SPELL, that learns long-range multimodal graphs to encode the inter-modal relationship between audio and visual data. We cast active speaker detection as a…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Sourya Roy , Kyle Min , Subarna Tripathi , Tanaya Guha , Somdeb Majumdar

Active speaker detection (ASD) systems are important modules for analyzing multi-talker conversations. They aim to detect which speakers or none are talking in a visual scene at any given time. Existing research on ASD does not agree on the…

Sound · Computer Science 2022-07-12 Abudukelimu Wuerkaixi , You Zhang , Zhiyao Duan , Changshui Zhang