English
Related papers

Related papers: SOAF: Scene Occlusion-aware Neural Acoustic Field

200 papers

Objects produce different sounds when hit, and humans can intuitively infer how an object might sound based on its appearance and material properties. Inspired by this intuition, we propose Visual Acoustic Fields, a framework that bridges…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Yuelei Li , Hyunjin Kim , Fangneng Zhan , Ri-Zhao Qiu , Mazeyu Ji , Xiaojun Shan , Xueyan Zou , Paul Liang , Hanspeter Pfister , Xiaolong Wang

In rapidly-evolving domains such as autonomous driving, the use of multiple sensors with different modalities is crucial to ensure high operational precision and stability. To correctly exploit the provided information by each sensor in a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Quentin Herau , Nathan Piasco , Moussab Bennehar , Luis Roldão , Dzmitry Tsishkou , Cyrille Migniot , Pascal Vasseur , Cédric Demonceaux

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

Multimedia · Computer Science 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman

Novel view synthesis of indoor scenes can be achieved by capturing a monocular video sequence of the environment. However, redundant information caused by artificial movements in the input video data reduces the efficiency of scene…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Zehao Wang , Han Zhou , Matthew B. Blaschko , Tinne Tuytelaars , Minye Wu

We propose a novel method for generating scene-aware training data for far-field automatic speech recognition. We use a deep learning-based estimator to non-intrusively compute the sub-band reverberation time of an environment from its…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-23 Zhenyu Tang , Dinesh Manocha

Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Huadai Liu , Tianyi Luo , Kaicheng Luo , Qikai Jiang , Peiwen Sun , Jialei Wang , Rongjie Huang , Qian Chen , Wen Wang , Xiangtai Li , Shiliang Zhang , Zhijie Yan , Zhou Zhao , Wei Xue

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

In this paper, we propose a new strategy for acoustic scene classification (ASC) , namely recognizing acoustic scenes through identifying distinct sound events. This differs from existing strategies, which focus on characterizing global…

Sound · Computer Science 2019-10-23 Hongwei Song , Jiqing Han , Shiwen Deng , Zhihao Du

Sight and hearing are two senses that play a vital role in human communication and scene understanding. To mimic human perception ability, audio-visual learning, aimed at developing computational approaches to learn from both audio and…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Yake Wei , Di Hu , Yapeng Tian , Xuelong Li

This paper presents a learning-based approach to synthesize the view from an arbitrary camera position given a sparse set of images. A key challenge for this novel view synthesis arises from the reconstruction process, when the views from…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Nan Meng , Kai Li , Jianzhuang Liu , Edmund Y. Lam

Structure from motion (SfM) enables us to reconstruct a scene via casual capture from cameras at different viewpoints, and novel view synthesis (NVS) allows us to render a captured scene from a new viewpoint. Both are hard with casual…

Identification and localization of sounds are both integral parts of computational auditory scene analysis. Although each can be solved separately, the goal of forming coherent auditory objects and achieving a comprehensive spatial scene…

Sound · Computer Science 2019-12-24 Ivo Trowitzsch , Christopher Schymura , Dorothea Kolossa , Klaus Obermayer

We present a method to synthesize novel views from a single $360^\circ$ panorama image based on the neural radiance field (NeRF). Prior studies in a similar setting rely on the neighborhood interpolation capability of multi-layer…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Shreyas Kulkarni , Peng Yin , Sebastian Scherer

We present a method for composing photorealistic scenes from captured images of objects. Our work builds upon neural radiance fields (NeRFs), which implicitly model the volumetric density and directionally-emitted radiance of a scene. While…

Computer Vision and Pattern Recognition · Computer Science 2020-12-16 Michelle Guo , Alireza Fathi , Jiajun Wu , Thomas Funkhouser

Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with…

This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Jiahe Li , Jiawei Zhang , Xiao Bai , Jun Zhou , Lin Gu

While experimentation with synthetic stimuli in abstracted listening situations has a long standing and successful history in hearing research, an increased interest exists on closing the remaining gap towards real-life listening by…

Accurate modeling of spatial acoustics is critical for immersive and intelligible audio in confined, resonant environments such as car cabins. Current tuning methods are manual, hardware-intensive, and static, failing to account for…

Sound · Computer Science 2025-10-10 Harshvardhan C. Takawale , Nirupam Roy , Phil Brown

Novel view synthesis refers to the problem of synthesizing novel viewpoints of a scene given the images from a few viewpoints. This is a fundamental problem in computer vision and graphics, and enables a vast variety of applications such as…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Nagabhushan Somraj
‹ Prev 1 3 4 5 6 7 10 Next ›