English
Related papers

Related papers: Audio Spatially-Guided Fusion for Audio-Visual Nav…

200 papers

Speaker diarization consists of assigning speech signals to people engaged in a dialogue. An audio-visual spatiotemporal diarization model is proposed. The model is well suited for challenging scenarios that consist of several participants…

Computer Vision and Pattern Recognition · Computer Science 2018-10-15 Israel D. Gebru , Silèye Ba , Xiaofei Li , Radu Horaud

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

Sound source localization is a typical and challenging task that predicts the location of sound sources in a video. Previous single-source methods mainly used the audio-visual association as clues to localize sounding objects in each image.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Shentong Mo , Yapeng Tian

Audio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interaction are required. The…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Shaofei Huang , Han Li , Yuqing Wang , Hongji Zhu , Jiao Dai , Jizhong Han , Wenge Rong , Si Liu

Vision to audition substitution devices are designed to convey visual information through auditory input. The acceptance of such systems depends heavily on their ease of use, training time, reliability and on the amount of coverage of…

Human-Computer Interaction · Computer Science 2020-10-20 Louis Commère , Sean U. N. Wood , Jean Rouat

There exists an unequivocal distinction between the sound produced by a static source and that produced by a moving one, especially when the source moves towards or away from the microphone. In this paper, we propose to use this connection…

Sound · Computer Science 2022-11-01 Moitreya Chatterjee , Narendra Ahuja , Anoop Cherian

Vision-Language Navigation requires the agent to follow natural language instructions to reach a specific target. The large discrepancy between seen and unseen environments makes it challenging for the agent to generalize well. Previous…

Computer Vision and Pattern Recognition · Computer Science 2022-09-13 Yujie Lu , Huiliang Zhang , Ping Nie , Weixi Feng , Wenda Xu , Xin Eric Wang , William Yang Wang

Recent years have witnessed the increasing application of place recognition in various environments, such as city roads, large buildings, and a mix of indoor and outdoor places. This task, however, still remains challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-24 Haowen Lai , Peng Yin , Sebastian Scherer

Audio-visual segmentation (AVS) is a challenging task that involves accurately segmenting sounding objects based on audio-visual cues. The effectiveness of audio-visual learning critically depends on achieving accurate cross-modal alignment…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Yuanhong Chen , Yuyuan Liu , Hu Wang , Fengbei Liu , Chong Wang , Helen Frazer , Gustavo Carneiro

Leveraging the synergy of both audio data and visual data is essential for understanding human emotions and behaviors, especially in in-the-wild setting. Traditional methods for integrating such multimodal information often stumble, leading…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Jun Yu , Zerui Zhang , Zhihong Wei , Gongpeng Zhao , Zhongpeng Cai , Yongqi Wang , Guochen Xie , Jichao Zhu , Wangyuan Zhu

In this paper, we present a novel deep fusion architecture for audio classification tasks. The multi-channel model presented is formed using deep convolution layers where different acoustic features are passed through each channel. To…

Sound · Computer Science 2018-11-05 Gaurav Bhatt , Akshita Gupta , Aditya Arora , Balasubramanian Raman

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving sources. First, we…

Sound · Computer Science 2025-09-19 Arvind Krishna Sridhar , Yinyi Guo , Erik Visser

We study the task of learning association between faces and voices, which is gaining interest in the multimodal community lately. These methods suffer from the deliberate crafting of negative mining procedures as well as the reliance on the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Abdul Hannan , Muhammad Arslan Manzoor , Shah Nawaz , Muhammad Irzam Liaqat , Markus Schedl , Mubashir Noman

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuta Kaneko , Abu Saleh Musa Miah , Najmul Hassan , Hyoun-Sup Lee , Si-Woong Jang , Jungpil Shin

Vision-language navigation is the task of directing an embodied agent to navigate in 3D scenes with natural language instructions. For the agent, inferring the long-term navigation target from visual-linguistic clues is crucial for reliable…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Yusheng Zhao , Jinyu Chen , Chen Gao , Wenguan Wang , Lirong Yang , Haibing Ren , Huaxia Xia , Si Liu

Audio-visual video segmentation (AVVS) aims to generate pixel-level maps of sound-producing objects that accurately align with the corresponding audio. However, existing methods often face temporal misalignment, where audio cues and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Kexin Li , Zongxin Yang , Yi Yang , Jun Xiao

An estimated 253 million people have visual impairments. These visual impairments affect everyday lives, and limit their understanding of the outside world. This can pose a risk to health from falling or collisions. We propose a solution to…

Human-Computer Interaction · Computer Science 2023-03-30 Alexander Mehta , Ritik Jalisatgi

The novelty of this study consists in a multi-modality approach to scene classification, where image and audio complement each other in a process of deep late fusion. The approach is demonstrated on a difficult classification problem,…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Jordan J. Bird , Diego R. Faria , Cristiano Premebida , Anikó Ekárt , George Vogiatzis

Audio tokenization has emerged as a critical component in end-to-end audio language models, enabling efficient discrete representation learning for both audio understanding and generation tasks. However, existing audio tokenizers face…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Xiangyu Zhang , Benjamin John Southwell , Siqi Pan , Xinlei Niu , Beena Ahmed , Julien Epps
‹ Prev 1 8 9 10 Next ›