English
Related papers

Related papers: AV-Gaze: A Study on the Effectiveness of Audio Gui…

200 papers

Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals. Existing audio-visual diarization datasets are mainly focused on indoor environments like meeting rooms or news studios, which are…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Eric Zhongcong Xu , Zeyang Song , Satoshi Tsutsui , Chao Feng , Mang Ye , Mike Zheng Shou

Many previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse musical video data, voice activity detection is a necessary…

Sound · Computer Science 2021-06-23 Yuanbo Hou , Zhesong Yu , Xia Liang , Xingjian Du , Bilei Zhu , Zejun Ma , Dick Botteldooren

Eye movements can reveal valuable insights into various aspects of human mental processes, physical well-being, and actions. Recently, several datasets have been made available that simultaneously record EEG activity and eye movements. This…

Signal Processing · Electrical Eng. & Systems 2023-08-14 Nina Weng , Martyna Plomecka , Manuel Kaufmann , Ard Kastrati , Roger Wattenhofer , Nicolas Langer

In recent years, the integration of vision and language understanding has led to significant advancements in artificial intelligence, particularly through Vision-Language Models (VLMs). However, existing VLMs face challenges in handling…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Kun Yan , Lei Ji , Zeyu Wang , Yuntao Wang , Nan Duan , Shuai Ma

We introduce Audio-Visual Affordance Grounding (AV-AG), a new task that segments object interaction regions from action sounds. Unlike existing approaches that rely on textual instructions or demonstration videos, which often limited by…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Lidong Lu , Guo Chen , Zhu Wei , Yicheng Liu , Tong Lu

Appearance-based gaze estimation from RGB images provides relatively unconstrained gaze tracking. We have previously proposed a gaze decomposition method that decomposes the gaze angle into the sum of a subject-independent gaze estimate…

Computer Vision and Pattern Recognition · Computer Science 2022-02-15 Zhaokang Chen , Bertram E. Shi

Affective computing research traditionally focused on labeling a person's emotion as one of a discrete number of classes e.g. happy or sad. In recent times, more attention has been given to continuous affect prediction across dimensions in…

Human-Computer Interaction · Computer Science 2018-03-06 Jonny O'Dwyer , Ronan Flynn , Niall Murray

Visually grounded speech models learn from images paired with spoken captions. By tagging images with soft text labels using a trained visual classifier with a fixed vocabulary, previous work has shown that it is possible to train a model…

Computation and Language · Computer Science 2021-06-24 Kayode Olaleye , Herman Kamper

Speech is understood better by using visual context; for this reason, there have been many attempts to use images to adapt automatic speech recognition (ASR) systems. Current work, however, has shown that visually adapted ASR models only…

Computation and Language · Computer Science 2020-02-19 Tejas Srinivasan , Ramon Sanabria , Florian Metze

Emotion recognition,as a step toward mind reading,seeks to infer internal states from external cues.Most existing methods rely on explicit signals-such as facial expressions,speech,or gestures-that reflect only bodily responses and overlook…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Mengke Song , Yuge Xie , Qi Cui , Luming Li , Xinyu Liu , Guotao Wang , Chenglizhao Chen , Shanchen Pang

In recent years, face recognition systems have achieved exceptional success due to promising advances in deep learning architectures. However, they still fail to achieve expected accuracy when matching profile images against a gallery of…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Moktari Mostofa , Mohammad Saeed Ebrahimi Saadabadi , Sahar Rahimi Malakshan , Nasser M. Nasrabadi

Effective assisted living environments must be able to perform inferences on how their occupants interact with one another as well as with surrounding objects. To accomplish this goal using a vision-based automated approach, multiple tasks…

Computer Vision and Pattern Recognition · Computer Science 2019-09-23 Philipe A. Dias , Damiano Malafronte , Henry Medeiros , Francesca Odone

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…

Computer Vision and Pattern Recognition · Computer Science 2018-06-14 Aviv Gabbay , Asaph Shamir , Shmuel Peleg

Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio…

Sound · Computer Science 2021-05-14 Xinyuan Qian , Maulik Madhavi , Zexu Pan , Jiadong Wang , Haizhou Li

The goal of our work is to use visual attention to enhance autonomous driving performance. We present two methods of predicting visual attention maps. The first method is a supervised learning approach in which we collect eye-gaze data for…

Computer Vision and Pattern Recognition · Computer Science 2018-12-06 Sourav Pal , Tharun Mohandoss , Pabitra Mitra

Obtaining demographics information from video is valuable for a range of real-world applications. While approaches that leverage facial features for gender inference are very successful in restrained environments, they do not work in most…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Andy Catruna , Adrian Cosma , Ion Emilian Radoi

Appearance-based gaze estimation (AGE) has achieved remarkable performance in constrained settings, yet we reveal a significant generalization gap where existing AGE models often fail in practical, unconstrained scenarios, particularly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Zhenhao Li , Zheng Liu , Seunghyun Lee , Amin Fadaeinejad , Yuanhao Yu

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this work, inspired by…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Xubo Liu , Qiushi Huang , Xinhao Mei , Haohe Liu , Qiuqiang Kong , Jianyuan Sun , Shengchen Li , Tom Ko , Yu Zhang , Lilian H. Tang , Mark D. Plumbley , Volkan Kılıç , Wenwu Wang

In a recent paper, we presented the KU Leuven audiovisual, gaze-controlled auditory attention decoding (AV-GC-AAD) dataset, in which we recorded electroencephalography (EEG) signals of participants attending to one out of two competing…

Signal Processing · Electrical Eng. & Systems 2024-12-03 Simon Geirnaert , Iustina Rotaru , Tom Francart , Alexander Bertrand

Augmented reality (AR) is emerging in visual search tasks for increasingly immersive interactions with virtual objects. We propose an AR approach providing visual and audio hints along with gaze-assisted instant post-task feedback for…

Human-Computer Interaction · Computer Science 2023-11-15 Yuchong Zhang , Adam Nowak , Yueming Xuan , Andrzej Romanowski , Morten Fjeld
‹ Prev 1 4 5 6 7 8 10 Next ›