English
Related papers

Related papers: Synthetic Wave-Geometric Impulse Responses for Imp…

200 papers

Speech is a means of communication which relies on both audio and visual information. The absence of one modality can often lead to confusion or misinterpretation of information. In this paper we present an end-to-end temporal model capable…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-17 Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition (ASR) systems. In many applications, it is of interest to estimate WER given a pair of a speech utterance and a…

Computation and Language · Computer Science 2024-04-29 Chanho Park , Mingjie Chen , Thomas Hain

This paper investigates a novel intelligent reflecting surface (IRS)-based symbiotic radio (SR) system architecture consisting of a transmitter, an IRS, and an information receiver (IR). The primary transmitter communicates with the IR and…

Information Theory · Computer Science 2022-02-09 Meng Hua , Qingqing Wu , Luxi Yang , Robert Schober , H. Vincent Poor

The scarcity of large-scale classroom speech data has hindered the development of AI-driven speech models for education. Classroom datasets remain limited and not publicly available, and the absence of dedicated classroom noise or Room…

Sound · Computer Science 2025-10-03 Ahmed Adel Attia , Jing Liu , Carol Espy Wilson

Ray tracing accelerated with graphics processing units (GPUs) is an accurate and efficient simulation technique of wireless communication channels. In this paper, we extend a GPU-accelerated ray tracer (RT) to support the effects of…

Signal Processing · Electrical Eng. & Systems 2024-02-21 Sara Sandh , Hamed Radpour , Benjamin Rainer , Markus Hofer , Thomas Zemen

In this work, we investigate application of generative speech enhancement to improve the robustness of ASR models in noisy and reverberant conditions. We employ a recently-proposed speech enhancement model based on Schr\"odinger bridge,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-09 Rauf Nasretdinov , Roman Korostik , Ante Jukić

In this study, we introduce a method for estimating sound fields in reverberant environments using a conditional invertible neural network (CINN). Sound field reconstruction can be hindered by experimental errors, limited spatial data,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-11 Xenofon Karakonstantis , Efren Fernandez-Grande , Peter Gerstoft

Benefiting from the development of deep learning, text-to-speech (TTS) techniques using clean speech have achieved significant performance improvements. The data collected from real scenes often contains noise and generally needs to be…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-06 Qiushi Zhu , Yu Gu , Rilin Chen , Chao Weng , Yuchen Hu , Lirong Dai , Jie Zhang

We investigate covert communication in an intelligent reflecting surface (IRS)-assisted symbiotic radio (SR) system under the parasitic SR (PSR) and the commensal SR (CSR) cases, where an IRS is exploited to create a double reflection link…

Information Theory · Computer Science 2024-11-18 Yunpeng Feng , Jian Chen , Lu Lv , Yuchen Zhou , Long Yang , Naofal Al-Dhahir , Fumiyuki Adachi

Synthetic data is widely used in speech recognition due to the availability of text-to-speech models, which facilitate adapting models to previously unseen text domains. However, existing methods suffer in performance when they fine-tune an…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Hsuan Su , Hua Farn , Fan-Yun Sun , Shang-Tse Chen , Hung-yi Lee

This paper presents dEchorate: a new database of measured multichannel Room Impulse Responses (RIRs) including annotations of early echo timings and 3D positions of microphones, real sources and image sources under different wall…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-28 Diego Di Carlo , Pinchas Tandeitnik , Cédric Foy , Antoine Deleforge , Nancy Bertin , Sharon Gannot

We present in this paper an informed single-channel dereverberation method based on conditional generation with diffusion models. With knowledge of the room impulse response, the anechoic utterance is generated via reverse diffusion using a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-22 Jean-Marie Lemercier , Simon Welker , Timo Gerkmann

One solution to automatic speech recognition (ASR) of overlapping speakers is to separate speech and then perform ASR on the separated signals. Commonly, the separator produces artefacts which often degrade ASR performance. Addressing this…

Modern neural network-based speech processing systems usually need to have reverberation resistance, so the training of such systems requires a large amount of reverberation data. In the process of system training, it is now more inclined…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-19 Dong Yang

We introduce a new cross-modal fusion technique designed for generative error correction in automatic speech recognition (ASR). Our methodology leverages both acoustic information and external linguistic representations to generate accurate…

Dereverberation is an important sub-task of Speech Enhancement (SE) to improve the signal's intelligibility and quality. However, it remains challenging because the reverberation is highly correlated with the signal. Furthermore, the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-05 Satvik Venkatesh , Philip Coleman , Arthur Benilov , Simon Brown , Selim Sheta , Frederic Roskam

Speech dereverberation aims to alleviate the negative impact of late reverberant reflections. The weighted prediction error (WPE) method is a well-established technique known for its superior performance in dereverberation. However, in…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-07 Ziye Yang , Mengfei Zhang , Jie Chen

We introduce HiFi-HARP, a large-scale dataset of 7th-order Higher-Order Ambisonic Room Impulse Responses (HOA-RIRs) consisting of more than 100,000 RIRs generated via a hybrid acoustic simulation in realistic indoor scenes. HiFi-HARP…

Sound · Computer Science 2025-10-27 Shivam Saini , Jürgen Peissig

The pseudo-periodicity of voiced speech can be exploited in several speech processing applications. This requires however that the precise locations of the Glottal Closure Instants (GCIs) are available. The focus of this paper is the…

Sound · Computer Science 2020-01-03 Thomas Drugman , Mark Thomas , Jon Gudnason , Patrick Naylor , Thierry Dutoit

While dynamic Neural Radiance Fields (NeRF) have shown success in high-fidelity 3D modeling of talking portraits, the slow training and inference speed severely obstruct their potential usage. In this paper, we propose an efficient…

Computer Vision and Pattern Recognition · Computer Science 2022-11-23 Jiaxiang Tang , Kaisiyuan Wang , Hang Zhou , Xiaokang Chen , Dongliang He , Tianshu Hu , Jingtuo Liu , Gang Zeng , Jingdong Wang