English
Related papers

Related papers: Audio Spatially-Guided Fusion for Audio-Visual Nav…

200 papers

Sound-guided object segmentation has drawn considerable attention for its potential to enhance multimodal perception. Previous methods primarily focus on developing advanced architectures to facilitate effective audio-visual interactions,…

Sound · Computer Science 2025-03-18 Chen Liu , Liying Yang , Peike Li , Dadong Wang , Lincheng Li , Xin Yu

Infrared and visible image fusion has garnered considerable attention owing to the strong complementarity of these two modalities in complex, harsh environments. While deep learning-based fusion methods have made remarkable advances in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Guihui Li , Bowei Dong , Kaizhi Dong , Jiayi Li , Haiyong Zheng

Automatic speech recognition (ASR) has reached a level of accuracy in recent years, that even outperforms humans in transcribing speech to text. Nevertheless, all current ASR approaches show a certain weakness against ambient noise. To…

Sound · Computer Science 2023-12-22 Christopher Simic , Tobias Bocklet

We study the problem of multimodal physical scene understanding, where an embodied agent needs to find fallen objects by inferring object properties, direction, and distance of an impact sound source. Previous works adopt feed-forward…

Robotics · Computer Science 2024-07-17 Jie Yin , Andrew Luo , Yilun Du , Anoop Cherian , Tim K. Marks , Jonathan Le Roux , Chuang Gan

Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 You Qin , Kai Liu , Shengqiong Wu , Kai Wang , Shijian Deng , Yapeng Tian , Junbin Xiao , Yazhou Xing , Yinghao Ma , Bobo Li , Roger Zimmermann , Lei Cui , Furu Wei , Jiebo Luo , Hao Fei

Geometric navigation is nowadays a well-established field of robotics and the research focus is shifting towards higher-level scene understanding, such as Semantic Mapping. When a robot needs to interact with its environment, it must be…

Robotics · Computer Science 2023-11-23 Federico Rollo , Gennaro Raiola , Andrea Zunino , Nikolaos Tsagarakis , Arash Ajoudani

We consider the problem of object goal navigation in unseen environments. Solving this problem requires learning of contextual semantic priors, a challenging endeavour given the spatial and semantic variability of indoor environments.…

Computer Vision and Pattern Recognition · Computer Science 2022-03-10 Georgios Georgakis , Bernadette Bucher , Karl Schmeckpeper , Siddharth Singh , Kostas Daniilidis

We introduce a novel deep learning-based audio-visual quality (AVQ) prediction model that leverages internal features from state-of-the-art unimodal predictors. Unlike prior approaches that rely on simple fusion strategies, our model…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Ina Salaj , Arijit Biswas

This work focuses on object goal visual navigation, aiming at finding the location of an object from a given class, where in each step the agent is provided with an egocentric RGB image of the scene. We propose to learn the agent's policy…

Computer Vision and Pattern Recognition · Computer Science 2021-04-21 Bar Mayo , Tamir Hazan , Ayellet Tal

Real-time open-vocabulary scene understanding is essential for efficient 3D perception in applications such as vision-language navigation, embodied intelligence, and augmented reality. However, existing methods suffer from imprecise…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Xiaofeng Jin , Matteo Frosi , Matteo Matteucci

The use of Autonomous Surface Vessels (ASVs) is growing rapidly. For safe and efficient surface auto-driving, a reliable perception system is crucial. Such systems allow the vessels to sense their surroundings and make decisions based on…

Robotics · Computer Science 2023-10-03 Xueyao Liang , Hu Xu , Yuwei Cheng

Audio-visual speech recognition (AVSR) can effectively and significantly improve the recognition rates of small-vocabulary systems, compared to their audio-only counterparts. For large-vocabulary systems, however, there are still many…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-13 Wentao Yu , Steffen Zeiler , Dorothea Kolossa

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audio-visual data…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Chuang Gan , Hang Zhao , Peihao Chen , David Cox , Antonio Torralba

Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Lionel Pibre , Francisco Madrigal , Cyrille Equoy , Frédéric Lerasle , Thomas Pellegrini , Julien Pinquier , Isabelle Ferrané

Two less addressed issues of deep reinforcement learning are (1) lack of generalization capability to new target goals, and (2) data inefficiency i.e., the model requires several (and often costly) episodes of trial and error to converge,…

Computer Vision and Pattern Recognition · Computer Science 2016-09-19 Yuke Zhu , Roozbeh Mottaghi , Eric Kolve , Joseph J. Lim , Abhinav Gupta , Li Fei-Fei , Ali Farhadi

Audio zooming, a signal processing technique, enables selective focusing and enhancement of sound signals from a specified region, attenuating others. While traditional beamforming and neural beamforming techniques, centered on creating a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-23 Meng Yu , Dong Yu

Recent embodied navigation approaches leveraging Vision-Language Models (VLMs) demonstrate strong generalization in versatile Vision-Language Navigation (VLN). However, reliable path planning in complex environments remains challenging due…

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-reading video…

Computer Vision and Pattern Recognition · Computer Science 2018-03-14 Matthijs Van keirsbilck , Bert Moons , Marian Verhelst

While video-to-audio generation has achieved remarkable progress in semantic and temporal alignment, most existing studies focus solely on these aspects, paying limited attention to the spatial perception and immersive quality of the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Yanan Wang , Linjie Ren , Zihao Li , Junyi Wang , Tian Gan

Image-goal navigation aims to steer an agent towards the goal location specified by an image. Most prior methods tackle this task by learning a navigation policy, which extracts visual features of goal and observation images, compares their…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Pengna Li , Kangyi Wu , Jingwen Fu , Sanping Zhou