English
Related papers

Related papers: Synthetic Wave-Geometric Impulse Responses for Imp…

200 papers

This paper introduces a novel data-driven strategy for synthesizing gramophone noise audio textures. A diffusion probabilistic model is applied to generate highly realistic quasiperiodic noises. The proposed model is designed to generate…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-01 Eloi Moliner , Vesa Välimäki

Room impulse responses (RIRs) are essential for many acoustic signal processing tasks, yet measuring them densely across space is often impractical. In this work, we propose RIR-Former, a grid-free, one-step feed-forward model for RIR…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Shaoheng Xu , Chunyi Sun , Jihui Zhang , Prasanga N. Samarasinghe , Thushara D. Abhayapala

Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model room-acoustic effects such as reverberation or room impulse responses (RIRs), and thus offer…

Sound · Computer Science 2026-05-04 Akira Takahashi , Ryosuke Sawata , Shusuke Takahashi , Yuki Mitsufuji

We synthesize both optical RGB and synthetic aperture radar (SAR) remote sensing images from land cover maps and auxiliary raster data using generative adversarial networks (GANs). In remote sensing, many types of data, such as digital…

Computer Vision and Pattern Recognition · Computer Science 2021-05-26 Gerald Baier , Antonin Deschemps , Michael Schmitt , Naoto Yokoya

We propose a novel method for generating scene-aware training data for far-field automatic speech recognition. We use a deep learning-based estimator to non-intrusively compute the sub-band reverberation time of an environment from its…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-23 Zhenyu Tang , Dinesh Manocha

Silent speech recognition (SSR) is a technology that recognizes speech content from non-acoustic speech-related biosignals. This paper utilizes an attention-enhanced temporal convolutional network architecture for contactless IR-UWB…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-01 Sunghwa Lee , Jaewon Yu

This paper proposes an efficient attempt to noisy speech emotion recognition (NSER). Conventional NSER approaches have proven effective in mitigating the impact of artificial noise sources, such as white Gaussian noise, but are limited to…

Sound · Computer Science 2026-01-13 Xiaohan Shi , Jiajun He , Xingfeng Li , Tomoki Toda

Recent breakthroughs in multi-talker ASR (MT-ASR) and speaker diarization (SD) rely on synthetic data to mitigate the scarcity of large-scale conversational recordings, yet the impact of specific simulation choices remains poorly…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-18 Alexander Polok , Ivan Medennikov , Jan Černocký , Shinji Watanabe , Lukáš Burget , Samuele Cornell

This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1) and Improving Speaker Distance Estimation (Task 2). We first…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-22 Christopher Ick , Gordon Wichern , Yoshiki Masuyama , François G. Germain , Jonathan Le Roux

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic information, and speaker…

We present a neural-network-based fast diffuse room impulse response generator (FAST-RIR) for generating room impulse responses (RIRs) for a given acoustic environment. Our FAST-RIR takes rectangular room dimensions, listener and speaker…

Sound · Computer Science 2022-02-08 Anton Ratnarajah , Shi-Xiong Zhang , Meng Yu , Zhenyu Tang , Dinesh Manocha , Dong Yu

We present a self-supervised speech restoration method without paired speech corpora. Because the previous general speech restoration method uses artificial paired data created by applying various distortions to high-quality speech corpora,…

In this paper, we propose to utilise diffusion models for data augmentation in speech emotion recognition (SER). In particular, we present an effective approach to utilise improved denoising diffusion probabilistic models (IDDPM) to…

Sound · Computer Science 2023-05-22 Ibrahim Malik , Siddique Latif , Raja Jurdak , Björn Schuller

We investigate the use of generative adversarial networks (GANs) in speech dereverberation for robust speech recognition. GANs have been recently studied for speech enhancement to remove additive noises, but there still lacks of a work to…

Sound · Computer Science 2019-01-01 Ke Wang , Junbo Zhang , Sining Sun , Yujun Wang , Fei Xiang , Lei Xie

Reconfigurable intelligent surfaces (RISs) have recently received widespread attention in the field of wireless communication. An RIS can be controlled to reflect incident waves from the transmitter towards the receiver; a feature that is…

Information Theory · Computer Science 2021-09-15 Jiangfeng Hu , Haifan Yin , Emil Björnson

Radio interferometry invariably suffers from an incomplete coverage of the spatial Fourier space, which leads to imaging artifacts. The current state-of-the-art technique is to create an image by Fourier-transforming the incomplete…

Instrumentation and Methods for Astrophysics · Physics 2024-12-19 F. Geyer , K. Schmidt , J. Kummer , M. Brüggen , H. W. Edler , D. Elsässer , F. Griese , A. Poggenpohl , L. Rustige , W. Rhode

Reverberant speech, denoting the speech signal degraded by reverberation, contains crucial knowledge of both anechoic source speech and room impulse response (RIR). This work proposes a variational Bayesian inference (VBI) framework with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-10 Pengyu Wang , Ying Fang , Xiaofei Li

Many students struggle with math word problems (MWPs), often finding it difficult to identify key information and select the appropriate mathematical operations. Schema-based instruction (SBI) is an evidence-based strategy that helps…

Machine Learning · Computer Science 2024-11-12 Prakhar Dixit , Tim Oates

Effective extraction and application of linguistic features are central to the enhancement of spoken Language IDentification (LID) performance. With the success of recent large models, such as GPT and Whisper, the potential to leverage such…

Computation and Language · Computer Science 2023-12-19 Peng Shen , Xuguang Lu , Hisashi Kawai

Whispering is a distinct form of speech known for its soft, breathy, and hushed characteristics, often used for private communication. The acoustic characteristics of whispered speech differ substantially from normally phonated speech and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-08 Zhaofeng Lin , Tanvina Patel , Odette Scharenborg
‹ Prev 1 4 5 6 7 8 10 Next ›