English
Related papers

Related papers: Pay Self-Attention to Audio-Visual Navigation

200 papers

The goal of Automatic Voice Over (AVO) is to generate speech in sync with a silent video given its text script. Recent AVO frameworks built upon text-to-speech synthesis (TTS) have shown impressive results. However, the current AVO learning…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-30 Junchen Lu , Berrak Sisman , Mingyang Zhang , Haizhou Li

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Arsha Nagrani , Shan Yang , Anurag Arnab , Aren Jansen , Cordelia Schmid , Chen Sun

A complex visual navigation task puts an agent in different situations which call for a diverse range of visual perception abilities. For example, to "go to the nearest chair", the agent might need to identify a chair in a living room using…

Computer Vision and Pattern Recognition · Computer Science 2021-08-05 Bokui Shen , Danfei Xu , Yuke Zhu , Leonidas J. Guibas , Li Fei-Fei , Silvio Savarese

We present a novel approach for image-goal navigation, where an agent navigates with a goal image rather than accurate target information, which is more challenging. Our goal is to decouple the learning of navigation goal planning,…

Robotics · Computer Science 2022-02-23 Qiaoyun Wu , Jun Wang , Jing Liang , Xiaoxi Gong , Dinesh Manocha

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-21 Abhinav Shukla , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

We propose a light-weight, self-supervised adaptation for a visual navigation agent to generalize to unseen environment. Given an embodied agent trained in a noiseless environment, our objective is to transfer the agent to a noisy…

Computer Vision and Pattern Recognition · Computer Science 2021-10-15 Eun Sun Lee , Junho Kim , Young Min Kim

We propose a cross attention transformer based method for multimodal sensor fusion to build a birds eye view of a vessels surroundings supporting safer autonomous marine navigation. The model deeply fuses multiview RGB and long wave…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Dimitrios Dagdilelis , Panagiotis Grigoriadis , Roberto Galeazzi

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sudha Krishnamurthy

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment…

Computer Vision and Pattern Recognition · Computer Science 2020-11-05 Soo-Whan Chung , Joon Son Chung , Hong-Goo Kang

The Vision-and-Language Navigation (VLN) task entails an agent following navigational instruction in photo-realistic unknown environments. This challenging task demands that the agent be aware of which instruction was completed, which…

Artificial Intelligence · Computer Science 2019-01-11 Chih-Yao Ma , Jiasen Lu , Zuxuan Wu , Ghassan AlRegib , Zsolt Kira , Richard Socher , Caiming Xiong

Vision-Language Navigation (VLN) is a core challenge in embodied AI, requiring agents to navigate real-world environments using natural language instructions. Current language model-based navigation systems operate on discrete topological…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Zhangyang Qi , Zhixiong Zhang , Yizhou Yu , Jiaqi Wang , Hengshuang Zhao

Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous, undifferentiated feedback. We present AMAVA, a novel real-time video-to-audio…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Benjamin Klein , Kazi Ruslan Rahman , Sanchita Ghose

Automatic emotion recognition (ER) has recently gained lot of interest due to its potential in many real-world applications. In this context, multimodal approaches have been shown to improve performance (over unimodal approaches) by…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 R Gnana Praveen , Eric Granger , Patrick Cardinal

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-03 Hiroshi Sato , Tsubasa Ochiai , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Shoko Araki

In this paper we propose a new framework - MoViLan (Modular Vision and Language) for execution of visually grounded natural language instructions for day to day indoor household tasks. While several data-driven, end-to-end learning…

Computer Vision and Pattern Recognition · Computer Science 2021-01-21 Homagni Saha , Fateme Fotouhif , Qisai Liu , Soumik Sarkar

Collaborative navigation becomes essential in situations of occluded scenarios in autonomous driving where independent driving policies are likely to lead to collisions. One promising approach to address this issue is through the use of…

Robotics · Computer Science 2024-12-12 Leandro Parada , Hanlin Tian , Jose Escribano , Panagiotis Angeloudis

Visual monitoring operations underwater require both observing the objects of interest in close-proximity, and tracking the few feature-rich areas necessary for state estimation.This paper introduces the first navigation framework, called…

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

Sound · Computer Science 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform additive fusion…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Jinbae Seo , Hyeongjun Kwon , Kwonyoung Kim , Jiyoung Lee , Kwanghoon Sohn

We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visual modalities while maintaining a compact model…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Mahrukh Awan , Asmar Nadeem , Armin Mustafa