中文
相关论文

相关论文: Scene-wide Acoustic Parameter Estimation

200 篇论文

This paper introduces a new framework for supervised sound source localization referred to as virtually-supervised learning. An acoustic shoe-box room simulator is used to generate a large number of binaural single-source audio scenes.…

声音 · 计算机科学 2017-03-21 Saurabh Kataria , Clément Gaultier , Antoine Deleforge

In recent years, dynamic parameterization of acoustic environments has garnered attention in audio processing. This focus includes room volume and reverberation time (RT60), which define local acoustics independent of sound source and…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Chunxi Wang , Maoshen Jia , Meiran Li , Changchun Bao , Wenyu Jin

We address the problem of estimating room impulse responses (RIRs) in noisy, uncontrolled environments where non-stationary sounds such as speech or footsteps corrupt conventional deconvolution. We propose AnyRIR, a non-intrusive method…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Kyung Yun Lee , Nils Meyer-Kahlen , Karolina Prawda , Vesa Välimäki , Sebastian J. Schlecht

The inference of the absorption configuration of an existing room solely using acoustic signals can be challenging. This research presents two methods for estimating the room dimensions and frequency-dependent absorption coefficients using…

声音 · 计算机科学 2023-04-26 Yuanxin Xia , Cheol-Ho Jeong

The recently proposed audio-visual scene-aware dialog task paves the way to a more data-driven way of learning virtual assistants, smart speakers and car navigation systems. However, very little is known to date about how to effectively…

计算机视觉与模式识别 · 计算机科学 2019-04-12 Idan Schwartz , Alexander Schwing , Tamir Hazan

Inverse rendering aims to estimate physical attributes of a scene, e.g., reflectance, geometry, and lighting, from image(s). Inverse rendering has been studied primarily for single objects or with methods that solve for only one of the…

计算机视觉与模式识别 · 计算机科学 2019-09-17 Soumyadip Sengupta , Jinwei Gu , Kihwan Kim , Guilin Liu , David W. Jacobs , Jan Kautz

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

音频与语音处理 · 电气工程与系统科学 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

Having knowledge on the room acoustic properties, e.g., the location of acoustic reflectors, allows to better reproduce the sound field as intended. Current state-of-the-art methods for room boundary detection using microphone measurements…

音频与语音处理 · 电气工程与系统科学 2022-06-09 Ellen Riemens , Pablo Martínez-Nuevo , Jorge Martinez , Martin Møller , Richard C. Hendriks

This paper investigates the application of environmental feature representations for room verification tasks and acoustic meta-data estimation. Audio recordings contain both speaker and non-speaker information. We refer to the…

声音 · 计算机科学 2022-03-10 Desmond Caulley

This paper introduces a novel image-based rendering technique for jointly estimating indoor lighting and thermal conditions from paired indoor-outdoor high dynamic range (HDR) panoramas. Our method uses the indoor panorama to estimate the…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Guanzhou Ji , Sriram Narayanan , Azadeh Sawyer , Srinivasa Narasimhan

We investigate an method for quantifying city characteristics based on impressions of a sound environment. The quantification of the city characteristics will be beneficial to government policy planning, tourism projects, etc. In this…

声音 · 计算机科学 2022-09-12 Yusuke Ono , Sunao Hara , Masanobu Abe

Rendering immersive spatial audio in virtual reality (VR) and video games demands a fast and accurate generation of room impulse responses (RIRs) to recreate auditory environments plausibly. However, the conventional methods for simulating…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Jackie Lin , Georg Götz , Sebastian J. Schlecht

Room impulse responses (RIRs) are fundamental to audio data augmentation, acoustic signal processing, and immersive audio rendering. While geometric simulators such as the image source method (ISM) can efficiently generate early…

音频与语音处理 · 电气工程与系统科学 2026-03-17 Zeyu Xu , Andreas Brendel , Albert G. Prinn , Emanuël A. P. Habets

We address the new problem of language-guided semantic style transfer of 3D indoor scenes. The input is a 3D indoor scene mesh and several phrases that describe the target scene. Firstly, 3D vertex coordinates are mapped to RGB residues by…

计算机视觉与模式识别 · 计算机科学 2022-08-17 Bu Jin , Beiwen Tian , Hao Zhao , Guyue Zhou

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Geometric acoustics is an efficient framework for room acoustics modeling, governed by the canonical time-dependent rendering equation. Acoustic radiance transfer (ART) solves the equation by discretization, modeling time- and…

声音 · 计算机科学 2026-04-17 Sungho Lee , Matteo Scerbo , Seungu Han , Min Jun Choi , Kyogu Lee , Enzo De Sena

We present an algorithm that fully reverses the shoebox image source method (ISM), a popular and widely used room impulse response (RIR) simulator for cuboid rooms introduced by Allen and Berkley in 1979. More precisely, given a discrete…

声音 · 计算机科学 2025-03-11 Tom Sprunck , Antoine Deleforge , Yannick Privat , Cédric Foy

For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images. We address this challenge by introducing Hypersim, a photorealistic synthetic dataset for holistic…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Mike Roberts , Jason Ramapuram , Anurag Ranjan , Atulit Kumar , Miguel Angel Bautista , Nathan Paczan , Russ Webb , Joshua M. Susskind

The Image Source Method (ISM) is one of the most employed techniques to calculate acoustic Room Impulse Responses (RIRs), however, its computational complexity grows fast with the reverberation time of the room and its computation time can…

音频与语音处理 · 电气工程与系统科学 2020-10-12 David Diaz-Guerra , Antonio Miguel , Jose R. Beltran

We propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, a SVBRDF, and 3D spatially-varying lighting. Because multi-view images provide a variety of information about the scene,…

计算机视觉与模式识别 · 计算机科学 2023-03-28 JunYong Choi , SeokYeong Lee , Haesol Park , Seung-Won Jung , Ig-Jae Kim , Junghyun Cho