English
Related papers

Related papers: FAST-RIR: Fast neural diffuse room impulse respons…

200 papers

This paper introduces a shoebox room simulator able to systematically generate synthetic datasets of binaural room impulse responses (BRIRs) given an arbitrary set of head-related transfer functions (HRTFs). The evaluation of machine…

Sound · Computer Science 2021-06-25 Roberto Barumerli , Daniele Bianchi , Michele Geronazzo , Federico Avanzini

We investigate the impact of more realistic room simulation for training far-field keyword spotting systems without fine-tuning on in-domain data. To this end, we study the impact of incorporating the following factors in the room impulse…

Sound · Computer Science 2020-11-19 Eric Bezzam , Robin Scheibler , Cyril Cadoux , Thibault Gisselbrecht

Rendering dynamic reverberation in a complicated acoustic space for moving sources and listeners is challenging but crucial for enhancing user immersion in extended-reality (XR) applications. Capturing spatially varying room impulse…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-10 Orchisama Das , Gloria Dal Santo , Sebastian J. Schlecht , Vesa Valimaki , Zoran Cvetkovic

State-of-the-art deep-learning-based voice activity detectors (VADs) are often trained with anechoic data. However, real acoustic environments are generally reverberant, which causes the performance to significantly deteriorate. To mitigate…

Sound · Computer Science 2021-06-28 Amir Ivry , Israel Cohen , Baruch Berdugo

We introduce Whisper-RIR-Mega, a benchmark dataset of paired clean and reverberant speech for evaluating automatic speech recognition (ASR) robustness to room acoustics. Each sample pairs a clean LibriSpeech utterance with the same…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Mandip Goswami

Room impulse responses (RIRs) are essential for many acoustic signal processing tasks, yet measuring them densely across space is often impractical. In this work, we propose RIR-Former, a grid-free, one-step feed-forward model for RIR…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Shaoheng Xu , Chunyi Sun , Jihui Zhang , Prasanga N. Samarasinghe , Thushara D. Abhayapala

Rings like gold, thuds like wood! The sound we hear in a scene is shaped not only by the spatial layout of the environment but also by the materials of the objects and surfaces within it. For instance, a room with wooden walls will produce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Mahnoor Fatima Saad , Sagnik Majumder , Kristen Grauman , Ziad Al-Halah

We present an efficient and realistic geometric acoustic simulation approach for generating and augmenting training data in speech-related machine learning tasks. Our physically-based acoustic simulation method is capable of modeling…

Sound · Computer Science 2021-09-28 Zhenyu Tang , Lianwu Chen , Bo Wu , Dong Yu , Dinesh Manocha

Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability and degraded performance under unseen conditions. Moreover,…

Sound · Computer Science 2026-02-11 Jackie Lin , Jiaqi Su , Nishit Anand , Zeyu Jin , Minje Kim , Paris Smaragdis

Fast Automatic Speech Recognition (ASR) is critical for latency-sensitive applications such as real-time captioning and meeting transcription. However, truly parallel ASR decoding remains challenging due to the sequential nature of…

Changes in room acoustics, such as modifications to surface absorption or the insertion of a scattering object, significantly impact measured room impulse responses (RIRs). These changes can affect the performance of systems used in echo…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Karolina Prawda

Neural Radiance Fields (NeRF) offer significant promise for generating photorealistic images and videos. However, existing mainstream neural rendering models often fall short in meeting the demands for immediacy and power efficiency in…

Hardware Architecture · Computer Science 2025-08-05 Fangxin Liu , Haomin Li , Bowen Zhu , Zongwu Wang , Zhuoran Song , Habing Guan , Li Jiang

Ensuring performance robustness for a variety of situations that can occur in real-world environments is one of the challenging tasks in sound event classification. One of the unpredictable and detrimental factors in performance, especially…

Sound · Computer Science 2021-04-22 Jaejun Lee , Donmoon Lee , Hyeong-Seok Choi , Kyogu Lee

Room acoustics analysis plays a central role in architectural design, audio engineering, speech intelligibility assessment, and hearing research. Despite the availability of standardized metrics such as reverberation time, clarity, and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-16 Mandip Goswami

We propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone…

Computation and Language · Computer Science 2015-09-02 Andreas Schwarz , Christian Huemmer , Roland Maas , Walter Kellermann

This paper focuses on room fingerprinting, a task involving the analysis of an audio recording to determine the specific volume and shape of the room in which it was captured. While it is relatively straightforward to determine the basic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-06 Jacob Bitterman , Daniel Levi , Hilel Hagai Diamandi , Sharon Gannot , Tal Rosenwein

The ability to accurately estimate room impulse responses (RIRs) is integral to many applications of spatial audio processing. Regrettably, estimating the RIR using ambient signals, such as speech or music, remains a challenging problem due…

Signal Processing · Electrical Eng. & Systems 2024-03-07 David Sundström , Anton Björkman , Andreas Jakobsson , Filip Elvander

A method is presented for estimating and reconstructing the sound field within a room using physics-informed neural networks. By incorporating a limited set of experimental room impulse responses as training data, this approach combines…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-03 Xenofon Karakonstantis , Diego Caviedes-Nozal , Antoine Richard , Efren Fernandez-Grande

Realistic audio synthesis that captures accurate acoustic phenomena is essential for creating immersive experiences in virtual and augmented reality. Synthesizing the sound received at any position relies on the estimation of impulse…

Sound · Computer Science 2024-11-12 Zitong Lan , Chenhao Zheng , Zhiwei Zheng , Mingmin Zhao

A deep learning approach has been widely applied in sequence modeling problems. In terms of automatic speech recognition (ASR), its performance has significantly been improved by increasing large speech corpus and deeper neural network.…

Computation and Language · Computer Science 2016-12-28 Zewang Zhang , Zheng Sun , Jiaqi Liu , Jingwen Chen , Zhao Huo , Xiao Zhang