English
Related papers

Related papers: The VVAD-LRS3 Dataset for Visual Voice Activity De…

200 papers

In this paper, we study the associations between human faces and voices. Audiovisual integration, specifically the integration of facial and vocal information is a well-researched area in neuroscience. It is shown that the overlapping…

Computer Vision and Pattern Recognition · Computer Science 2018-11-05 Changil Kim , Hijung Valentina Shin , Tae-Hyun Oh , Alexandre Kaspar , Mohamed Elgharib , Wojciech Matusik

Audio-visual speaker recognition is one of the tasks in the recent 2019 NIST speaker recognition evaluation (SRE). Studies in neuroscience and computer science all point to the fact that vision and auditory neural signals interact in the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Ruijie Tao , Rohan Kumar Das , Haizhou Li

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

Sound · Computer Science 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie

Recently, the remarkable success of ChatGPT has sparked a renewed wave of interest in artificial intelligence (AI), and the advancements in visual language models (VLMs) have pushed this enthusiasm to new heights. Differring from previous…

Artificial Intelligence · Computer Science 2025-01-03 Lijie Tao , Haokui Zhang , Haizhao Jing , Yu Liu , Dawei Yan , Guoting Wei , Xizhe Xue

Language has always been one of humanity's defining characteristics. Visual Language Identification (VLI) is a relatively new field of research that is complex and largely understudied. In this paper, we present a preliminary study in which…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Lucia Cascone , Michele Nappi , Fabio Narducci

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-13 Otavio Braga , Olivier Siohan

Speech activity detection (SAD) plays an important role in current speech processing systems, including automatic speech recognition (ASR). SAD is particularly difficult in environments with acoustic noise. A practical solution is to…

Computation and Language · Computer Science 2023-05-15 Fei Tao , Carlos Busso

As research on neural volumetric video reconstruction and compression flourishes, there is a need for diverse and realistic datasets, which can be used to develop and validate reconstruction and compression models. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Adrian Azzarelli , Ge Gao , Ho Man Kwan , Fan Zhang , Nantheera Anantrasirichai , Ollie Moolan-Feroze , David Bull

Spatio-temporal contexts are crucial in understanding human actions in videos. Recent state-of-the-art Convolutional Neural Network (ConvNet) based action recognition systems frequently involve 3D spatio-temporal ConvNet filters, chunking…

Computer Vision and Pattern Recognition · Computer Science 2018-05-09 Yunfeng Wang , Wengang Zhou , Qilin Zhang , Xiaotian Zhu , Houqiang Li

Social intelligence, the ability to interpret emotions, intentions, and behaviors, is essential for effective communication and adaptive responses. As robots and AI systems become more prevalent in caregiving, healthcare, and education, the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Erika Mori , Yue Qiu , Hirokatsu Kataoka , Yoshimitsu Aoki

Recent advances in detecting arbitrary objects in the real world are trained and evaluated on object detection datasets with a relatively restricted vocabulary. To facilitate the development of more general visual object detection, we…

Computer Vision and Pattern Recognition · Computer Science 2023-10-06 Jiaqi Wang , Pan Zhang , Tao Chu , Yuhang Cao , Yujie Zhou , Tong Wu , Bin Wang , Conghui He , Dahua Lin

Recent advances in deep learning based large vocabulary con- tinuous speech recognition (LVCSR) invoke growing demands in large scale speech transcription. The inference process of a speech recognizer is to find a sequence of labels whose…

Computation and Language · Computer Science 2018-08-03 Zhehuai Chen

Audio-visual recognition (AVR) has been considered as a solution for speech recognition tasks when the audio is corrupted, as well as a visual recognition method used for speaker verification in multi-speaker scenarios. The approach of AVR…

Computer Vision and Pattern Recognition · Computer Science 2017-11-01 Amirsina Torfi , Seyed Mehdi Iranmanesh , Nasser M. Nasrabadi , Jeremy Dawson

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-25 Alkis Koudounas , Moreno La Quatra , Marco Sabato Siniscalchi , Elena Baralis

Teleconference or telepresence based on virtual reality (VR) headmount display (HMD) device is a very interesting and promising application since HMD can provide immersive feelings for users. However, in order to facilitate face-to-face…

Computer Vision and Pattern Recognition · Computer Science 2019-01-23 Guoxian Song , Jianfei Cai , Tat-Jen Cham , Jianmin Zheng , Juyong Zhang , Henry Fuchs

Open-Vocabulary Segmentation (OVS) methods offer promising capabilities in detecting unseen object categories, but the category must be known and needs to be provided by a human, either via a text prompt or pre-labeled datasets, thus…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Weijie Wei , Osman Ülger , Fatemeh Karimi Nejadasl , Theo Gevers , Martin R. Oswald

Lipreading is the task of decoding text from the movement of a speaker's mouth. Traditional approaches separated the problem into two stages: designing or learning visual features, and prediction. More recent deep lipreading approaches are…

Machine Learning · Computer Science 2016-12-19 Yannis M. Assael , Brendan Shillingford , Shimon Whiteson , Nando de Freitas

Audio-Visual Speech Recognition (AVSR) integrates acoustic and visual information to enhance robustness in adverse acoustic conditions. Recent advances in Large Language Models (LLMs) have yielded competitive automatic speech recognition…

Sound · Computer Science 2026-03-05 Fei Su , Cancan Li , Juan Liu , Wei Ju , Hongbin Suo , Ming Li

A core process in human cognition is analogical mapping: the ability to identify a similar relational structure between different situations. We introduce a novel task, Visual Analogies of Situation Recognition, adapting the classical…

Computer Vision and Pattern Recognition · Computer Science 2022-12-12 Yonatan Bitton , Ron Yosef , Eli Strugo , Dafna Shahaf , Roy Schwartz , Gabriel Stanovsky

We present a minimalistic but effective neural network that computes dense facial correspondences in highly unconstrained RGB images. Our network learns a per-pixel flow and a matchability mask between 2D input photographs of a person and…

Computer Vision and Pattern Recognition · Computer Science 2017-09-05 Ronald Yu , Shunsuke Saito , Haoxiang Li , Duygu Ceylan , Hao Li