English
Related papers

Related papers: SonoWorld: From One Image to a 3D Audio-Visual Sce…

200 papers

Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a…

Sound · Computer Science 2025-03-18 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Jiankang Deng , Xiatian Zhu

We infer and generate three-dimensional (3D) scene information from a single input image and without supervision. This problem is under-explored, with most prior work relying on supervision from, e.g., 3D ground-truth, multiple images of a…

Computer Vision and Pattern Recognition · Computer Science 2020-04-20 Sai Rajeswar , Fahim Mannan , Florian Golemo , Jérôme Parent-Lévesque , David Vazquez , Derek Nowrouzezahrai , Aaron Courville

Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with stereo supervision.…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Hang Zhou , Xudong Xu , Dahua Lin , Xiaogang Wang , Ziwei Liu

We present Worldsheet, a method for novel view synthesis using just a single RGB image as input. The main insight is that simply shrink-wrapping a planar mesh sheet onto the input image, consistent with the learned intermediate depth,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Ronghang Hu , Nikhila Ravi , Alexander C. Berg , Deepak Pathak

In audio-visual navigation, an agent intelligently travels through a complex, unmapped 3D environment using both sights and sounds to find a sound source (e.g., a phone ringing in another room). Existing models learn to act at a fixed…

Computer Vision and Pattern Recognition · Computer Science 2021-02-12 Changan Chen , Sagnik Majumder , Ziad Al-Halah , Ruohan Gao , Santhosh Kumar Ramakrishnan , Kristen Grauman

One major goal of vision is to infer physical models of objects, surfaces, and their layout from sensors. In this paper, we aim to interpret indoor scenes from one RGBD image. Our representation encodes the layout of orthogonal walls and…

Computer Vision and Pattern Recognition · Computer Science 2018-11-15 Chuhang Zou , Ruiqi Guo , Zhizhong Li , Derek Hoiem

This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise. Our approach is based on generative inverse sampling, where we model clean speech and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-03 Yochai Yemini , Yoav Ellinson , Rami Ben-Ari , Sharon Gannot , Ethan Fetaya

Neural volumetric representations have become a widely adopted model for radiance fields in 3D scenes. These representations are fully implicit or hybrid function approximators of the instantaneous volumetric radiance in a scene, which are…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Yuval Bahat , Yuxuan Zhang , Hendrik Sommerhoff , Andreas Kolb , Felix Heide

Current open-vocabulary scene graph generation algorithms highly rely on both 3D scene point cloud data and posed RGB-D images and thus have limited applications in scenarios where RGB-D images or camera poses are not readily available. To…

Robotics · Computer Science 2024-09-17 Yifan Xu , Ziming Luo , Qianwei Wang , Vineet Kamat , Carol Menassa

Deep generative models allow for photorealistic image synthesis at high resolutions. But for many applications, this is not enough: content creation also needs to be controllable. While several recent works investigate how to disentangle…

Computer Vision and Pattern Recognition · Computer Science 2021-04-30 Michael Niemeyer , Andreas Geiger

We explore the generation of visualisations of audio latent spaces using an audio-to-image generation pipeline. We believe this can help with the interpretability of audio latent spaces. We demonstrate a variety of results on the NSynth…

Sound · Computer Science 2022-12-08 Nicolas Jonason , Bob L. T. Sturm

Creating high-fidelity 3D models of indoor environments is essential for applications in design, virtual reality, and robotics. However, manual 3D modeling remains time-consuming and labor-intensive. While recent advances in generative AI…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Chuan Fang , Heng Li , Yixun Liang , Jia Zheng , Yongsen Mao , Yuan Liu , Rui Tang , Zihan Zhou , Ping Tan

We report Zero123++, an image-conditioned diffusion model for generating 3D-consistent multi-view images from a single input view. To take full advantage of pretrained 2D generative priors, we develop various conditioning and training…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Ruoxi Shi , Hansheng Chen , Zhuoyang Zhang , Minghua Liu , Chao Xu , Xinyue Wei , Linghao Chen , Chong Zeng , Hao Su

There has been a growing interest in the task of generating sound for silent videos, primarily because of its practicality in streamlining video post-production. However, existing methods for video-sound generation attempt to directly…

Multimedia · Computer Science 2024-04-04 Zhifeng Xie , Shengye Yu , Qile He , Mengtian Li

We introduce and explore a new multimodal input representation for vision-language models: acoustic field video. Unlike conventional video (RGB with stereo/mono audio), our video stream provides a spatially grounded visualization of sound…

Human-Computer Interaction · Computer Science 2026-01-27 Daehwa Kim , Chris Harrison

Reconstructing detailed 3D scenes from single-view images remains a challenging task due to limitations in existing approaches, which primarily focus on geometric shape recovery, overlooking object appearances and fine shape details. To…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Yixin Chen , Junfeng Ni , Nan Jiang , Yaowei Zhang , Yixin Zhu , Siyuan Huang

We consider the challenging problem of audio to animated video generation. We propose a novel method OneShotAu2AV to generate an animated video of arbitrary length using an audio clip and a single unseen image of a person as an input. The…

Computer Vision and Pattern Recognition · Computer Science 2021-02-22 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall , Mujtaba Hasan , Pranshu Agarwal , Dipankar Sarkar

We present a method for learning to generate unbounded flythrough videos of natural scenes starting from a single view, where this capability is learned from a collection of single photographs, without requiring camera poses or even…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Zhengqi Li , Qianqian Wang , Noah Snavely , Angjoo Kanazawa

Generating large-scale 3D scenes cannot simply apply existing 3D object synthesis technique since 3D scenes usually hold complex spatial configurations and consist of a number of objects at varying scales. We thus propose a practical and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Qihang Zhang , Yinghao Xu , Yujun Shen , Bo Dai , Bolei Zhou , Ceyuan Yang

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Suzhen Wang , Lincheng Li , Yu Ding , Changjie Fan , Xin Yu
‹ Prev 1 8 9 10 Next ›