English
Related papers

Related papers: Dynamical Audio-Visual Navigation: Catching Unhear…

200 papers

We have observed significant progress in visual navigation for embodied agents. A common assumption in studying visual navigation is that the environments are static; this is a limiting assumption. Intelligent navigation may involve…

Computer Vision and Pattern Recognition · Computer Science 2021-05-04 Kuo-Hao Zeng , Luca Weihs , Ali Farhadi , Roozbeh Mottaghi

Autonomous navigation in unknown environments requires multi-scale spatial understanding that captures geometric details, topological connectivity, and global structure to support high-level decision making under partial observability.…

Robotics · Computer Science 2026-04-22 Kuankuan Sima , Longbin Tang , Zhenyu Yang , Haozhe Ma , Lin Zhao

Autonomous underwater navigation remains a challenging problem due to limited sensing capabilities and the difficulty of constructing accurate maps in underwater environments. In this paper, we propose a Diffusion-based Underwater Visual…

Robotics · Computer Science 2025-09-04 Jinghe Yang , Minh-Quan Le , Mingming Gong , Ye Pu

We present a novel method for populating 3D indoor scenes with virtual humans that can navigate in the environment and interact with objects in a realistic manner. Existing approaches rely on training sequences that contain captured human…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Kaifeng Zhao , Yan Zhang , Shaofei Wang , Thabo Beeler , Siyu Tang

Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task -- real-world audio-visual scene synthesis -- and a…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Susan Liang , Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Multi-agent reinforcement learning has been used as an effective means to study emergent communication between agents, yet little focus has been given to continuous acoustic communication. This would be more akin to human language…

Computation and Language · Computer Science 2023-05-03 Kevin Eloff , Okko Räsänen , Herman A. Engelbrecht , Arnu Pretorius , Herman Kamper

Sound sources localization using multichannel signal processing has been a subject of active research for decades. In recent years, the use of deep learning in audio signal processing has allowed to drastically improve performances for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Hadrien Pujol , Éric Bavu , Alexandre Garcia

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds…

Multimedia · Computer Science 2024-04-02 Siva Sai Nagender Vasireddy , Chenxu Zhang , Xiaohu Guo , Yapeng Tian

Safe visual navigation is critical for indoor mobile robots operating in cluttered environments. Existing benchmarks, however, often neglect collisions or are designed for outdoor scenarios, making them unsuitable for indoor visual…

Robotics · Computer Science 2026-03-05 Jaewon Lee , Jaeseok Heo , Gunmin Lee , Howoong Jun , Jeongwoo Oh , Songhwai Oh

Speech activity detection (SAD) plays an important role in current speech processing systems, including automatic speech recognition (ASR). SAD is particularly difficult in environments with acoustic noise. A practical solution is to…

Computation and Language · Computer Science 2023-05-15 Fei Tao , Carlos Busso

It is common to implicitly assume access to intelligently captured inputs (e.g., photos from a human photographer), yet autonomously capturing good observations is itself a major challenge. We address the problem of learning to look around:…

Computer Vision and Pattern Recognition · Computer Science 2017-12-22 Dinesh Jayaraman , Kristen Grauman

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D Visual…

Machine Learning · Computer Science 2025-08-14 Duc Cao-Dinh , Khai Le-Duc , Anh Dao , Bach Phan Tat , Chris Ngo , Duy M. H. Nguyen , Nguyen X. Khanh , Thanh Nguyen-Tang

Audio-Visual Localization (AVL) aims to identify sound-emitting sources within a visual scene. However, existing studies focus on image-level audio-visual associations, failing to capture temporal dynamics. Moreover, they assume simplified…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Hahyeon Choi , Junhoo Lee , Nojun Kwak

Tomorrow's robots will need to distinguish useful information from noise when performing different tasks. A household robot for instance may continuously receive a plethora of information about the home, but needs to focus on just a small…

Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Luca Barbisan , Marco Levorato , Fabrizio Riente

This paper presents an autonomous navigation framework for reaching a goal in unknown 3D cluttered environments. The framework consists of three main components. First, a computationally efficient method for mapping the environment from the…

In the context of visual navigation, the capacity to map a novel environment is necessary for an agent to exploit its observation history in the considered place and efficiently reach known goals. This ability can be associated with spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-04-26 Pierre Marza , Laetitia Matignon , Olivier Simonin , Christian Wolf

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

Exploration is one of the core challenges in reinforcement learning. A common formulation of curiosity-driven exploration uses the difference between the real future and the future predicted by a learned model. However, predicting the…

Machine Learning · Computer Science 2021-01-19 Victoria Dean , Shubham Tulsiani , Abhinav Gupta