中文
相关论文

相关论文: Audio-Visual World Models: Towards Multisensory Im…

200 篇论文

Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous, undifferentiated feedback. We present AMAVA, a novel real-time video-to-audio…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Benjamin Klein , Kazi Ruslan Rahman , Sanchita Ghose

World Models have emerged as a powerful paradigm for learning compact, predictive representations of environment dynamics, enabling agents to reason, plan, and generalize beyond direct experience. Despite recent interest in World Models,…

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

We explore active audio-visual separation for dynamic sound sources, where an embodied agent moves intelligently in a 3D environment to continuously isolate the time-varying audio stream being emitted by an object of interest. The agent…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Sagnik Majumder , Kristen Grauman

Audio-Visual Large Language Models (AVLLMs) are emerging as unified interfaces to multimodal perception. We present the first mechanistic interpretability study of AVLLMs, analyzing how audio and visual features evolve and fuse through…

人工智能 · 计算机科学 2026-04-06 Ramaneswaran Selvakumar , Kaousheik Jayakumar , S Sakshi , Sreyan Ghosh , Ruohan Gao , Dinesh Manocha

World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego-vehicle perspective, leaving the infrastructure viewpoint unexplored. We argue that…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Siyuan Meng , Chengbo Ai

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

声音 · 计算机科学 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Predicting future sensory states is crucial for learning agents such as robots, drones, and autonomous vehicles. In this paper, we couple multiple sensory modalities with exploratory actions and propose a predictive neural network…

机器人学 · 计算机科学 2021-09-17 Xiaohui Chen , Ramtin Hosseini , Karen Panetta , Jivko Sinapov

Multimodal research and applications are becoming more commonplace as Virtual Reality (VR) technology integrates different sensory feedback, enabling the recreation of real spaces in an audio-visual context. Within VR experiences, numerous…

音频与语音处理 · 电气工程与系统科学 2025-04-08 Mauricio Flores-Vargas , Enda Bates , Rachel McDonnell

Audio-visual navigation of an agent towards locating an audio goal is a challenging task especially when the audio is sporadic or the environment is noisy. In this paper, we present CAVEN, a Conversation-based Audio-Visual Embodied…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Xiulong Liu , Sudipta Paul , Moitreya Chatterjee , Anoop Cherian

Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a…

声音 · 计算机科学 2025-03-18 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Jiankang Deng , Xiatian Zhu

There have been many attempts to build multimodal dialog systems that can respond to a question about given audio-visual information, and the representative task for such systems is the Audio Visual Scene-Aware Dialog (AVSD). Most…

计算与语言 · 计算机科学 2022-02-22 Yoshihiro Yamazaki , Shota Orihashi , Ryo Masumura , Mihiro Uchida , Akihiko Takashima

We introduce Diffusion World Model (DWM), a conditional diffusion model capable of predicting multistep future states and rewards concurrently. As opposed to traditional one-step dynamics models, DWM offers long-horizon predictions in a…

机器学习 · 计算机科学 2024-10-17 Zihan Ding , Amy Zhang , Yuandong Tian , Qinqing Zheng

Current visual generation methods can produce high quality videos guided by texts. However, effectively controlling object dynamics remains a challenge. This work explores audio as a cue to generate temporally synchronized image animations.…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Lin Zhang , Shentong Mo , Yijing Zhang , Pedro Morgado

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

Recent progress in 3D reconstruction has made it easy to create realistic digital twins from everyday environments. However, current digital twins remain largely static and are limited to navigation and view synthesis without embodied…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Byungjun Kim , Taeksoo Kim , Junyoung Lee , Hanbyul Joo

Robotic imitation learning has advanced from solving static tasks to addressing dynamic interaction scenarios, but testing and evaluation remain costly and challenging due to the need for real-time interaction with dynamic environments. We…

Understanding and replicating the real world is a critical challenge in Artificial General Intelligence (AGI) research. To achieve this, many existing approaches, such as world models, aim to capture the fundamental principles governing the…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Yuqi Hu , Longguang Wang , Xian Liu , Ling-Hao Chen , Yuwei Guo , Yukai Shi , Ce Liu , Anyi Rao , Zeyu Wang , Hui Xiong

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…