English
Related papers

Related papers: StoRIR: Stochastic Room Impulse Response Generatio…

200 papers

Stochastic resonance (SR) could amplify weak electric-field signals in nonlinear systems by means of the externally injected noises. Here we propose and experimentally demonstrate a modified SR method, termed squeezing-induced SR,…

This paper proposes an efficient attempt to noisy speech emotion recognition (NSER). Conventional NSER approaches have proven effective in mitigating the impact of artificial noise sources, such as white Gaussian noise, but are limited to…

Sound · Computer Science 2026-01-13 Xiaohan Shi , Jiajun He , Xingfeng Li , Tomoki Toda

Data augmentation has proven to be effective in training neural networks. Recently, a method called RandAug was proposed, randomly selecting data augmentation techniques from a predefined search space. RandAug has demonstrated significant…

An initial real-time speech enhancement method is presented to reduce the effects of additive noise. The method operates in the frequency domain and is a form of spectral subtraction. Initially, minimum statistics are used to generate an…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-22 Georgios Ioannides , Vasilios Rallis

Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated to the audio input. While previous approaches make crucial…

Sound · Computer Science 2022-04-29 Dan Oneata , Horia Cucu

Own voice pickup for hearables in noisy environments benefits from using both an outer and an in-ear microphone outside and inside the occluded ear. Due to environmental noise recorded at both microphones, and amplification of the own voice…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-20 Mattes Ohlenbusch , Christian Rollwage , Simon Doclo

We present an algorithm that fully reverses the shoebox image source method (ISM), a popular and widely used room impulse response (RIR) simulator for cuboid rooms introduced by Allen and Berkley in 1979. More precisely, given a discrete…

Sound · Computer Science 2025-03-11 Tom Sprunck , Antoine Deleforge , Yannick Privat , Cédric Foy

Automatic speech emotion recognition (SER) is a challenging task that plays a crucial role in natural human-computer interaction. One of the main challenges in SER is data scarcity, i.e., insufficient amounts of carefully labeled data to…

Sound · Computer Science 2021-08-17 Sarala Padi , Seyed Omid Sadjadi , Dinesh Manocha , Ram D. Sriram

This paper presents a Multi-Modal Environment-Aware Network (MEAN-RIR), which uses an encoder-decoder framework to predict room impulse response (RIR) based on multi-level environmental information from audio, visual, and textual sources.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-08 Jiajian Chen , Jiakang Chen , Hang Chen , Qing Wang , Yu Gao , Jun Du

Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability and degraded performance under unseen conditions. Moreover,…

Sound · Computer Science 2026-02-11 Jackie Lin , Jiaqi Su , Nishit Anand , Zeyu Jin , Minje Kim , Paris Smaragdis

Noise reduction techniques based on deep learning have demonstrated impressive performance in enhancing the overall quality of recorded speech. While these approaches are highly performant, their application in audio engineering can be…

Sound · Computer Science 2023-10-18 Christian J. Steinmetz , Thomas Walther , Joshua D. Reiss

In this paper, we present HOMULA-RIR, a dataset of room impulse responses (RIRs) acquired using both higher-order microphones (HOMs) and a uniform linear array (ULA), in order to model a remote attendance teleconferencing scenario.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-22 Federico Miotello , Paolo Ostan , Mirco Pezzoli , Luca Comanducci , Alberto Bernardini , Fabio Antonacci , Augusto Sarti

The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advancements, existing models often fail to achieve a reliable…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-25 Jaemin Jung , Jaehun Kim , Inkyu Shin , Joon Son Chung

As humans, we hear sound every second of our life. The sound we hear is often affected by the acoustics of the environment surrounding us. For example, a spacious hall leads to more reverberation. Room Impulse Responses (RIR) are commonly…

Artificial Intelligence · Computer Science 2023-10-10 Yinfeng Yu , Changan Chen , Lele Cao , Fangkai Yang , Fuchun Sun

Latest advances in deep spatial filtering for Ambisonics demonstrate strong performance in stationary multi-speaker scenarios by rotating the sound field toward a target speaker prior to multi-channel enhancement. For applicability in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-22 Jakob Kienegger , Timo Gerkmann

Self-induced stochastic resonance (SISR) is the emergence of coherent oscillations in slow-fast excitable systems driven solely by noise, without external periodic forcing or proximity to a bifurcation. This work presents a physics-informed…

Machine Learning · Computer Science 2026-01-29 Divyesh Savaliya , Marius E. Yamakou

Stochastic resonance is a non-linear phenomenon, in which the sensitivity of signal detectors can be enhanced by adding random noise to the detector input. Here, we demonstrate that noise can also improve the information flux in recurrent…

Neurons and Cognition · Quantitative Biology 2018-11-30 Patrick Krauss , Karin Prebeck , Achim Schilling , Claus Metzner

Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is underexplored. We investigate the use of a pretraining stage based…

Speech recognition has of late become a practical technology for real world applications. Aiming at speech-driven text retrieval, which facilitates retrieving information with spoken queries, we propose a method to integrate speech…

Computation and Language · Computer Science 2007-05-23 Atsushi Fujii , Katunobu Itou , Tetsuya Ishikawa

This paper introduces a novel method for RGB-Guided Resolution Enhancement of infrared (IR) images called Guided IR Resolution Enhancement (GIRRE). In the area of single image super resolution (SISR) there exists a wide variety of…

Image and Video Processing · Electrical Eng. & Systems 2023-09-13 Marcel Trammer , Nils Genser , Jürgen Seiler