English
Related papers

Related papers: Multiple-Speaker Localization Based on Direct-Path…

200 papers

In this paper we address the problem of simultaneously tracking several moving audio sources, namely the problem of estimating source trajectories from a sequence of observed features. We propose to use the von Mises distribution to model…

Sound · Computer Science 2019-04-11 Yutong Ban , Xavier Alameda-PIneda , Christine Evers , Radu Horaud

Speech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or…

Sound · Computer Science 2025-07-14 Cheng Chi , Xiaoyu Li , Yuxuan Ke , Qunping Ni , Yao Ge , Xiaodong Li , Chengshi Zheng

Binaural acoustic source localization is important to human listeners for spatial awareness, communication and safety. In this paper, an end-to-end binaural localization model for speech in noise is presented. A lightweight convolutional…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-29 Vikas Tokala , Eric Grinstein , Rory Brooks , Mike Brookes , Simon Doclo , Jesper Jensen , Patrick A. Naylor

Gaussian process (GP) audio source separation is a time-domain approach that circumvents the inherent phase approximation issue of spectrogram based methods. Furthermore, through its kernel, GPs elegantly incorporate prior knowledge about…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-22 Pablo A. Alvarado , Mauricio A. Álvarez , Dan Stowell

Neural Text-to-Speech (TTS) systems find broad applications in voice assistants, e-learning, and audiobook creation. The pursuit of modern models, like Diffusion Models (DMs), holds promise for achieving high-fidelity, real-time speech…

Sound · Computer Science 2024-04-02 Xiang Li , Fan Bu , Ambuj Mehrish , Yingting Li , Jiale Han , Bo Cheng , Soujanya Poria

Conventional audio equalization is a static process that requires manual and cumbersome adjustments to adapt to changing listening contexts (e.g., mood, location, or social setting). In this paper, we introduce a Large Language Model…

In hearing aid applications, an important objective is to accurately estimate the direction of arrival (DOA) of multiple speakers in noisy and reverberant environments. Recently, we proposed a binaural DOA estimation method, where the DOAs…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-11 Daniel Fejgin , Simon Doclo

Non-line-of-sight localization in signal-deprived environments is a challenging yet pertinent problem. Acoustic methods in such predominantly indoor scenarios encounter difficulty due to the reverberant nature. In this study, we aim to…

Machine Learning · Computer Science 2024-04-03 Yi Di Yuan , Swee Liang Wong , Jonathan Pan

We propose an advance Steered Response Power (SRP) method for localizing multiple sources. While conventional SRP performs well in adverse conditions, it remains to struggle in scenarios with closely neighboring sources, resulting in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-21 Wei-Ting Lai , Lachlan Birnie , Xingyu Chen , Amy Bastine , Thushara D. Abhayapala , Prasanga N. Samarasinghe

Speaker localization in a reverberant environment is a fundamental problem in audio signal processing. Many solutions have been developed to tackle this problem. However, previous algorithms typically assume a stationary environment in…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-29 Daniel A. Mitchell , Boaz Rafaely

In speaker verification, traditional models often emphasize modeling long-term contextual features to capture global speaker characteristics. However, this approach can neglect fine-grained voiceprint information, which contains highly…

Sound · Computer Science 2025-05-07 Ya Li , Bin Zhou , Bo Hu

Extracting direct-path spatial feature is crucial for sound source localization in adverse acoustic environments. This paper proposes the IPDnet, a neural network that estimates direct-path inter-channel phase difference (DP-IPD) of sound…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-14 Yabo Wang , Bing Yang , Xiaofei Li

Point set registration is an essential step in many computer vision applications, such as 3D reconstruction and SLAM. Although there exist many registration algorithms for different purposes, however, this topic is still challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2022-11-22 Jin Zhang , Mingyang Zhao , Xin Jiang , Dong-Ming Yan

This paper describes a sound source localization (SSL) technique that combines an $\alpha$-stable model for the observed signal with a neural network-based approach for modeling steering vectors. Specifically, a physics-informed neural…

We derive an asymptotic expansion for the log likelihood of Gaussian mixture models (GMMs) with equal covariance matrices in the low signal-to-noise regime. The expansion reveals an intimate connection between two types of algorithms for…

Statistics Theory · Mathematics 2020-06-30 Anya Katsevich , Afonso Bandeira

Recent data- and learning-based sound source localization (SSL) methods have shown strong performance in challenging acoustic scenarios. However, little work has been done on adapting such methods to track consistently multiple sources…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 David Diaz-Guerra , Archontis Politis , Tuomas Virtanen

Localizing linearly moving sound sources using microphone arrays is challenging as the transient nature of the signal leads to relatively short observation periods. Commonly, a moving focus is used and most methods operate at least…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-21 Christian H. Kasess , Wolfgang Kreuzer , Prateek Soni , Holger Waubke

Many computer vision problems can be posed as learning a low-dimensional subspace from high dimensional data. The low rank matrix factorization (LRMF) represents a commonly utilized subspace learning strategy. Most of the current LRMF…

Computer Vision and Pattern Recognition · Computer Science 2016-09-21 Xiangyong Cao , Qian Zhao , Deyu Meng , Yang Chen , Zongben Xu

Real-world measurements often comprise a dominant signal contaminated by a noisy background. Robustly estimating the dominant signal in practice has been a fundamental statistical problem. Classically, mixture models have been used to…

Computation · Statistics 2026-05-20 Ananyabrata Barua , Ayanendranath Basu

The images and sounds that we perceive undergo subtle but geometrically consistent changes as we rotate our heads. In this paper, we use these cues to solve a problem we call Sound Localization from Motion (SLfM): jointly estimating camera…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Ziyang Chen , Shengyi Qian , Andrew Owens
‹ Prev 1 8 9 10 Next ›