English
Related papers

Related papers: RevRIR: Joint Reverberant Speech and Room Impulse …

200 papers

Realistic sound simulation plays a critical role in many applications. A key element in sound simulation is the room impulse response (RIR), which characterizes how sound propagates from a source to a listener within a given space. Recent…

Sound · Computer Science 2025-09-19 Chen Si , Qianyi Wu , Chaitanya Amballa , Romit Roy Choudhury

This study addresses robust automatic speech recognition (ASR) by introducing a Conformer-based acoustic model. The proposed model builds on the wide residual bi-directional long short-term memory network (WRBN) with utterance-wise dropout…

Sound · Computer Science 2022-10-21 Yufeng Yang , Peidong Wang , DeLiang Wang

Speaker identification typically involves three stages. First, a front-end speaker embedding model is trained to embed utterance and speaker profiles. Second, a scoring function is applied between a runtime utterance and each speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-22 Zhenning Tan , Yuguang Yang , Eunjung Han , Andreas Stolcke

This paper investigates the application of environmental feature representations for room verification tasks and acoustic meta-data estimation. Audio recordings contain both speaker and non-speaker information. We refer to the…

Sound · Computer Science 2022-03-10 Desmond Caulley

Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-Language Multirate and Multimodal Representations (CALM), an…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-09 Vin Sachidananda , Shao-Yen Tseng , Erik Marchi , Sachin Kajarekar , Panayiotis Georgiou

The ASVspoof 2021 benchmark, a widely-used evaluation framework for anti-spoofing, consists of two subsets: Logical Access (LA) and Deepfake (DF), featuring samples with varied coding characteristics and compression artifacts. Notably, the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Hieu-Thi Luong , Duc-Tuan Truong , Kong Aik Lee , Eng Siong Chng

With the advent of modern AI architectures, a shift has happened towards end-to-end architectures. This pivot has led to neural architectures being trained without domain-specific biases/knowledge, optimized according to the task. We in…

Sound · Computer Science 2025-05-08 Prateek Verma

Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embedding space in which corresponding items from the two…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-22 Paul Primus , Florian Schmid , Gerhard Widmer

The image source method (ISM) is often used to simulate room acoustics due to its ease of use and computational efficiency. The standard ISM is limited to simulations of room impulse responses between point sources and omnidirectional…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-08 Zeyu Xu , Adrian Herzog , Alexander Lodermeyer , Emanuël A. P. Habets , Albert G. Prinn

In this paper, we propose a model to perform speech dereverberation by estimating its spectral magnitude from the reverberant counterpart. Our models are capable of extracting features that take into account both short and long-term…

Sound · Computer Science 2017-11-20 Joao Felipe Santos , Tiago H. Falk

We propose multi-microphone complex spectral mapping, a simple way of applying deep learning for time-varying non-linear beamforming, for speaker separation in reverberant conditions. We aim at both speaker separation and dereverberation.…

Sound · Computer Science 2021-05-25 Zhong-Qiu Wang , Peidong Wang , DeLiang Wang

Room geometry is important prior information for implementing realistic 3D audio rendering. For this reason, various room geometry inference (RGI) methods have been developed by utilizing the time-of-arrival (TOA) or…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-25 Inmo Yeon , Jung-Woo Choi

Room Impulse Response (RIR) prediction at arbitrary receiver positions is essential for practical applications such as spatial audio rendering. We propose Neural Acoustic Multipole Splatting (NAMS), which synthesizes RIRs at unseen receiver…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-03 Geonwoo Baek , Jung-Woo Choi

Binaural reproduction aims to deliver immersive spatial audio with high perceptual realism over headphones. Loss functions play a central role in optimizing and evaluating algorithms that generate binaural signals. However, traditional…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-03 Boaz Rafaely , Stefan Weinzierl , Or Berebi , Fabian Brinkmann

In multi-room environments, modelling the sound propagation is complex due to the coupling of rooms and diverse source-receiver positions. A common scenario is when the source and the receiver are in different rooms without a clear line of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-19 Kyung Yun Lee , Nils Meyer-Kahlen , Georg Götz , U. Peter Svensson , Sebastian J. Schlecht , Vesa Välimäki

Immersion in virtual and augmented reality solutions is reliant on plausible spatial audio. However, plausibly representing a space for immersive audio often requires many individual acoustic measurements of source-microphone pairs with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-07 Ben Heritage , Fiona Ryder , Michael McLoughlin , Karolina Prawda

Real-time magnetic resonance imaging (rtMRI) of speech production enables non-invasive visualization of dynamic vocal-tract motion and is valuable for speech science and clinical assessment. However, rtMRI is fundamentally constrained by…

Inverse rendering aims to estimate physical attributes of a scene, e.g., reflectance, geometry, and lighting, from image(s). Inverse rendering has been studied primarily for single objects or with methods that solve for only one of the…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Soumyadip Sengupta , Jinwei Gu , Kihwan Kim , Guilin Liu , David W. Jacobs , Jan Kautz

Radio Frequency Fingerprinting (RFF) using deep learning has gained attention as a complementary approach to cryptographic authentication, offering resistance to spoofing, replay attacks, and key leakage. While most RFF approaches rely on…

Signal Processing · Electrical Eng. & Systems 2026-05-18 Fawaz Abdul Razak , Yasin Yilmaz

We present an indoor acoustic simulation framework that supports both ultrasonic and audible signaling. The framework opens the opportunity for fast indoor acoustic data generation and positioning development. The improved…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-22 Daan Delabie , Chesney Buyle , Bert Cox , Liesbet Van der Perre , Lieven De Strycker