English
Related papers

Related papers: Binaural sound source localization using a hybrid …

200 papers

Foundation models (FMs), that are trained on broad data at scale and are adaptable to a wide range of downstream tasks, have brought large interest in the research community. Benefiting from the diverse data sources such as different…

Computation and Language · Computer Science 2023-02-06 Bo Li , Dongseong Hwang , Zhouyuan Huo , Junwen Bai , Guru Prakash , Tara N. Sainath , Khe Chai Sim , Yu Zhang , Wei Han , Trevor Strohman , Francoise Beaufays

In this work, a parameterized eigenvalue problem is analyzed for a phononic array in a 2D stress wave scattering setup, and a corresponding sensing application of this system is proposed to achieve source angle localization. The phononic…

Applied Physics · Physics 2021-11-02 Weidi Wang , Amir Ashkan Mokhtari , Ankit Srivastava , Alireza V. Amirkhizi

Human Activity Recognition (HAR) benefits various application domains, including health and elderly care. Traditional HAR involves constructing pipelines reliant on centralized user data, which can pose privacy concerns as they necessitate…

Precise elevation perception in binaural audio remains a challenge, despite extensive research on head-related transfer functions (HRTFs) and spectral cues. While prior studies have advanced our understanding of sound localization cues, the…

Signal Processing · Electrical Eng. & Systems 2025-03-17 Juan Antonio De Rus , Mario Montagud , Jesus Lopez-Ballester , Francesc J. Ferri , Maximo Cobos

In this work we apply Amplitude Modulation Spectrum (AMS) features to the source localization problem. Our approach computes 36 bilateral features for 2s long signal segments and estimates the azimuthal directions of a sound source through…

Sound · Computer Science 2018-12-07 Semih Ağcaer , Rainer Martin

Binaural acoustic source localization is important to human listeners for spatial awareness, communication and safety. In this paper, an end-to-end binaural localization model for speech in noise is presented. A lightweight convolutional…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-29 Vikas Tokala , Eric Grinstein , Rory Brooks , Mike Brookes , Simon Doclo , Jesper Jensen , Patrick A. Naylor

Subjective evaluations are critical for assessing the perceptual realism of sounds in audio-synthesis driven technologies like augmented and virtual reality. However, they are challenging to set up, fatiguing for users, and expensive. In…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-22 Pranay Manocha , Anurag Kumar , Buye Xu , Anjali Menon , Israel D. Gebru , Vamsi K. Ithapu , Paul Calamia

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

We present a transformer-based speech-declipping model that effectively recovers clipped signals across a wide range of input signal-to-distortion ratios (SDRs). While recent time-domain deep neural network (DNN)-based declippers have…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Younghoo Kwon , Jung-Woo Choi

While the community keeps promoting end-to-end models over conventional hybrid models, which usually are long short-term memory (LSTM) models trained with a cross entropy criterion followed by a sequence discriminative training criterion,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-18 Jinyu Li , Rui Zhao , Eric Sun , Jeremy H. M. Wong , Amit Das , Zhong Meng , Yifan Gong

Parametric sound field synthesis methods, such as the Spatial Decomposition Method (SDM) and Higher-Order Spatial Impulse Response Rendering (HO-SIRR), are widely used for the analysis and auralization of sound fields. This paper studies…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-04 Alan Pawlak , Hyunkook Lee , Aki Mäkivirta , Thomas Lund

Full-field ultra-high-speed (UHS) x-ray imaging experiments have been well established to characterize various processes and phenomena. However, the potential of UHS experiments through the joint acquisition of x-ray videos with distinct…

Image and Video Processing · Electrical Eng. & Systems 2024-11-28 Songyuan Tang , Tekin Bicer , Tao Sun , Kamel Fezzaa , Samuel J. Clark

In this paper we address the problems of modeling the acoustic space generated by a full-spectrum sound source and of using the learned model for the localization and separation of multiple sources that simultaneously emit sparse-spectrum…

Sound · Computer Science 2015-02-06 Antoine Deleforge , Florence Forbes , Radu Horaud

This paper introduces a hybrid computational framework for the multi-frequency inverse source problem governed by the Helmholtz equation. By integrating a classical Fourier method with a deep convolutional neural network, we address the…

Analysis of PDEs · Mathematics 2026-01-05 Hao Chen , Yan Chang , Yukun Guo , Yuliang Wang

We introduce bidirectional edge diffraction response function (BEDRF), a new approach to model wave diffraction around edges with path tracing. The diffraction part of the wave is expressed as an integration on path space, and the wave-edge…

Sound · Computer Science 2023-06-06 Chunxiao Cao , Zili An , Zhong Ren , Dinesh Manocha , Kun Zhou

Human-robot interaction in natural settings requires filtering out the different sources of sounds from the environment. Such ability usually involves the use of microphone arrays to localize, track and separate sound sources online.…

Audio and Speech Processing · Electrical Eng. & Systems 2018-12-04 Francois Grondin , Francois Michaud

Beamforming with desired directivity patterns using compact microphone arrays is essential in many audio applications. Directivity patterns achievable using traditional beamformers depend on the number of microphones and the array aperture.…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-24 Weilong Huang , Srikanth Raj Chetupalli , Mhd Modar Halimeh , Oliver Thiergart , Emanuël A. P. Habets

We propose a natural way to generalize relative transfer functions (RTFs) to more than one source. We first prove that such a generalization is not possible using a single multichannel spectro-temporal observation, regardless of the number…

Sound · Computer Science 2015-07-02 Antoine Deleforge , Sharon Gannot , Walter Kellermann

We propose a transfer learning framework for sound source reconstruction in Near-field Acoustic Holography (NAH), which adapts a well-trained data-driven model from one type of sound source to another using a physics-informed procedure. The…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-16 Xinmeng Luan , Mirco Pezzoli , Fabio Antonacci , Augusto Sarti

Ultrasound, alone or in concert with circulating microbubble contrast agents, has emerged as a promising modality for therapy and imaging of brain diseases. While this has become possible due to advancements in aberration correction…

Signal Processing · Electrical Eng. & Systems 2020-01-08 Scott Schoen , Costas D. Arvanitis