English
Related papers

Related papers: Towards Generalisable Audio Representations for Au…

200 papers

Speech enhancement is an essential task of improving speech quality in noise scenario. Several state-of-the-art approaches have introduced visual information for speech enhancement,since the visual aspect of speech is essentially unaffected…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Xinmeng Xu , Yang Wang , Dongxiang Xu , Yiyuan Peng , Cong Zhang , Jie Jia , Binbin Chen

The audio-visual speech fusion strategy AV Align has shown significant performance improvements in audio-visual speech recognition (AVSR) on the challenging LRS2 dataset. Performance improvements range between 7% and 30% depending on the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-20 George Sterpu , Christian Saam , Naomi Harte

Voice Conversion (VC) emerged as a significant domain of research in the field of speech synthesis in recent years due to its emerging application in voice-assisting technology, automated movie dubbing, and speech-to-singing conversion to…

Sound · Computer Science 2021-04-27 Sandipan Dhar , Nanda Dulal Jana , Swagatam Das

Vision and voice are two vital keys for agents' interaction and learning. In this paper, we present a novel indoor navigation model called Memory Vision-Voice Indoor Navigation (MVV-IN), which receives voice commands and analyzes multimodal…

Computer Vision and Pattern Recognition · Computer Science 2020-09-02 Liqi Yan , Dongfang Liu , Yaoxian Song , Changbin Yu

Autonomous underwater vehicles (AUVs) rely on a variety of sensors - acoustic, inertial and visual - for intelligent decision making. Due to its non-intrusive, passive nature, and high information content, vision is an attractive sensing…

Computer Vision and Pattern Recognition · Computer Science 2018-01-15 Cameron Fabbri , Md Jahidul Islam , Junaed Sattar

Vision-and-language navigation (VLN) agents are trained to navigate in real-world environments by following natural language instructions. A major challenge in VLN is the limited availability of training data, which hinders the models'…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Zi-Yi Dou , Feng Gao , Nanyun Peng

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions where traditional…

Sound · Computer Science 2025-05-16 Detao Bai , Zhiheng Ma , Xihan Wei , Liefeng Bo

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

Multimedia · Computer Science 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

Vision-and-language navigation (VLN) asks an agent to follow a given language instruction to navigate through a real 3D environment. Despite significant advances, conventional VLN agents are trained typically under disturbance-free…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Bingqian Lin , Yanxin Long , Yi Zhu , Fengda Zhu , Xiaodan Liang , Qixiang Ye , Liang Lin

In this work we use deep reinforcement learning to create an autonomous agent that can navigate in a two-dimensional space using only raw auditory sensory information from the environment, a problem that has received very little attention…

Sound · Computer Science 2021-05-17 Petros Giannakopoulos , Aggelos Pikrakis , Yannis Cotronis

The aim of audio-visual segmentation (AVS) is to precisely differentiate audible objects within videos down to the pixel level. Traditional approaches often tackle this challenge by combining information from various modalities, where the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Dawei Hao , Yuxin Mao , Bowen He , Xiaodong Han , Yuchao Dai , Yiran Zhong

Vision-and-Language Navigation (VLN) is a task where agents must decide how to move through a 3D environment to reach a goal by grounding natural language instructions to the visual surroundings. One of the problems of the VLN task is data…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Tsu-Jui Fu , Xin Eric Wang , Matthew Peterson , Scott Grafton , Miguel Eckstein , William Yang Wang

Vision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robots or outdoor…

Robotics · Computer Science 2025-12-23 Xu Liu , Yu Liu , Hanshuo Qiu , Yang Qirong , Zhouhui Lian

Self-supervised learning (SSL) approaches, such as contrastive and generative methods, have advanced environmental sound representation learning using unlabeled data. However, how these approaches can complement each other within a unified…

Sound · Computer Science 2025-10-29 Sivan Ding , Julia Wilkins , Magdalena Fuentes , Juan Pablo Bello

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

Face videos accompanied by audio have become integral to our daily lives, while they often suffer from complex degradations. Most face video restoration methods neglect the intrinsic correlations between the visual and audio features,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Yuqin Cao , Yixuan Gao , Wei Sun , Xiaohong Liu , Yulun Zhang , Xiongkuo Min

What is a good visual representation for autonomous agents? We address this question in the context of semantic visual navigation, which is the problem of a robot finding its way through a complex environment to a target object, e.g. go to…

Computer Vision and Pattern Recognition · Computer Science 2019-07-04 Arsalan Mousavian , Alexander Toshev , Marek Fiser , Jana Kosecka , Ayzaan Wahid , James Davidson

In the Vision-and-Language Navigation (VLN) field, agents are tasked with navigating real-world scenes guided by linguistic instructions. Enabling the agent to adhere to instructions throughout the process of navigation represents a…

Artificial Intelligence · Computer Science 2024-05-28 Wen Hanlin

This paper presents a novel approach for the Vision-and-Language Navigation (VLN) task in continuous 3D environments, which requires an autonomous agent to follow natural language instructions in unseen environments. Existing end-to-end…

‹ Prev 1 3 4 5 6 7 10 Next ›