English
Related papers

Related papers: Estimation of the Direct-Path Relative Transfer Fu…

200 papers

Transformer models are powerful sequence-to-sequence architectures that are capable of directly mapping speech inputs to transcriptions or translations. However, the mechanism for modeling positions in this model was tailored for text…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-21 Ngoc-Quan Pham , Thanh-Le Ha , Tuan-Nam Nguyen , Thai-Son Nguyen , Elizabeth Salesky , Sebastian Stueker , Jan Niehues , Alexander Waibel

Fundamental rate-distortion-perception (RDP) trade-offs arise in applications requiring maintained perceptual quality of reconstructed data, such as neural image compression. When compressed data is transmitted over public communication…

Information Theory · Computer Science 2026-04-23 Gustaf Åhlgren , Onur Günlü

This article presents a method for estimating and reconstructing the spatial energy distribution pattern of natural speech, which is crucial for achieving realistic vocal presence in virtual communication settings. The method comprises two…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-06 Camille Noufi , Dejan Markovic , Peter Dodds

Blind estimation of acoustic room parameters such as the reverberation time $T_\mathrm{60}$ and the direct-to-reverberation ratio ($\mathrm{DRR}$) is still a challenging task, especially in case of blind estimation from reverberant speech…

Sound · Computer Science 2015-10-16 Feifei Xiong , Stefan Goetze , Bernd T. Meyer

The second-order harmonic (2f) component generated by twin-rotary compressor is a dominant low-frequency noise source of variable refrigerant flow (VRF) outdoor units, yet its amplitude fluctuates strongly with environmental thermal load…

Signal Processing · Electrical Eng. & Systems 2026-05-05 ZhiWei Su , Ding Wang , Yuan Guo , Yang Qiao , HongJun Cao

Acoustic scene perception involves describing the type of sounds, their timing, their direction and distance, as well as their loudness and reverberation. While audio language models excel in sound recognition, single-channel input…

Sound · Computer Science 2025-10-08 Xilin Jiang , Hannes Gamper , Sebastian Braun

Non-interactive and linear experiences like cinema film offer high quality surround sound audio to enhance immersion, however the listener's experience is usually fixed to a single acoustic perspective. With the rise of virtual reality,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-30 Lachlan Birnie , Thushara Abhayapala , Vladimir Tourbabin , Prasanga Samarasinghe

Voice conversion is a common speech synthesis task which can be solved in different ways depending on a particular real-world scenario. The most challenging one often referred to as one-shot many-to-many voice conversion consists in copying…

Sound · Computer Science 2022-08-05 Vadim Popov , Ivan Vovk , Vladimir Gogoryan , Tasnima Sadekova , Mikhail Kudinov , Jiansheng Wei

Psychoacoustic experiments have shown that directional properties of the direct sound, salient reflections, and the late reverberation of an acoustic room response can have a distinct influence on the auditory perception of a given room.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-28 Thomas Deppisch , Sebastià V. Amengual Garí , Paul Calamia , Jens Ahrens

The state-of-the-art automotive radars employ multidimensional discrete Fourier transforms (DFT) in order to estimate various target parameters. The DFT is implemented using the fast Fourier transform (FFT), at sample and computational…

Signal Processing · Electrical Eng. & Systems 2018-01-16 Shaogang Wang , Vishal M. Patel , Athina Petropulu

We investigate the effectiveness of convolutive prediction, a novel formulation of linear prediction for speech dereverberation, for speaker separation in reverberant conditions. The key idea is to first use a deep neural network (DNN) to…

Sound · Computer Science 2021-08-17 Zhong-Qiu Wang , Gordon Wichern , Jonathan Le Roux

This paper presents a novel approach for indoor acoustic source localization using microphone arrays and based on a Convolutional Neural Network (CNN). The proposed solution is, to the best of our knowledge, the first published work in…

Sound · Computer Science 2019-02-01 Juan Manuel Vera-Diaz , Daniel Pizarro , Javier Macias-Guarasa

Estimation of the direction-of-arrival (DOA) of sound sources is an important step in sound field analysis. Rigid spherical microphone arrays allow the calculation of a compact spherical harmonic representation of the sound field. A basic…

Sound · Computer Science 2018-03-06 Mert Burkay Coteli , Orhun Olgun , Huseyin Hacihabiboglu

Personalized head-related transfer functions (HRTFs) are essential for ensuring a realistic auditory experience over headphones, because they take into account individual anatomical differences that affect listening. Most machine learning…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 You Zhang , Andrew Francl , Ruohan Gao , Paul Calamia , Zhiyao Duan , Ishwarya Ananthabhotla

Recently, the rectified flow (RF) has emerged as the new state-of-the-art among flow-based diffusion models due to its high efficiency advantage in straight path sampling, especially with the amazing images generated by a series of RF…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Zhiyuan Ma , Ruixun Liu , Sixian Liu , Jianjun Li , Bowen Zhou

We present a novel sound localization algorithm for a non-line-of-sight (NLOS) sound source in indoor environments. Our approach exploits the diffraction properties of sound waves as they bend around a barrier or an obstacle in the scene.…

Robotics · Computer Science 2018-09-21 Inkyu An , Doheon Lee , Jung-woo Choi , Dinesh Manocha , Sung-eui Yoon

Inverse problems, which involve estimating parameters from incomplete or noisy observations, arise in various fields such as medical imaging, geophysics, and signal processing. These problems are often ill-posed, requiring regularization…

Image and Video Processing · Electrical Eng. & Systems 2026-03-03 Shadab Ahamed , Eldad Haber

Speech dereverberation in distant-microphone scenarios remains challenging due to the high correlation between reverberation and target signals, often leading to poor generalization in real-world environments. We propose IF-CorrNet, a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Ui-Hyeop Shin , Jun Hyung Kim , Jangyeon Kim , Wooseok Kim , Hyung-Min Park

This paper proposes a blind estimation method based on the modulation transfer function and Schroeder model for estimating reverberation time in seven-octave bands. Therefore, the speech transmission index and five room-acoustic parameters…

Sound · Computer Science 2021-03-16 Suradej Duangpummet , Jessada Karnjana , Waree Kongprawechnon , Masashi Unoki

An interpolation method for region-to-region acoustic transfer functions (ATFs) based on kernel ridge regression with an adaptive kernel is proposed. Most current ATF interpolation methods do not incorporate the acoustic properties for…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-08 Juliano G. C. Ribeiro , Shoichi Koyama , Hiroshi Saruwatari