中文
相关论文

相关论文: Multiple Sound Source Localization with SVD-PHAT

200 篇论文

To phased microphone array for sound source localization, algorithm with both high computational efficiency and high precision is a persistent pursuit. In this paper convolutional neural network (CNN) a kind of deep learning is…

音频与语音处理 · 电气工程与系统科学 2018-02-14 Wei Ma , Xun Liu

Graph-based Transform (GT) has been recently leveraged successfully in the signal processing domain, specifically for compression purposes. In this paper, we employ the GBT, as well as the Singular Value Decomposition (SVD) with the goal to…

音频与语音处理 · 电气工程与系统科学 2020-03-19 Majid Farzaneh , Rahil Mahdian Toroghi

We propose a data-driven sparse recovery framework for hybrid spherical linear microphone arrays using singular value decomposition (SVD) of the transfer operator. The SVD yields orthogonal microphone and field modes, reducing to spherical…

音频与语音处理 · 电气工程与系统科学 2026-02-04 Shunxi Xu , Thushara Abhayapala , Craig T. Jin

High-resolution array detectors are widely used in single-particle tracking, but their performance is limited by excess noise from background light and dark current. As pixel resolution increases, the diminished signal per pixel exacerbates…

量子物理 · 物理学 2025-12-16 Chao-Ning Hu , Jun Xin , Xiao-Ming Lu

In this paper, we propose a convolutional recurrent neural network for joint sound event localization and detection (SELD) of multiple overlapping sound events in three-dimensional (3D) space. The proposed network takes a sequence of…

声音 · 计算机科学 2018-12-18 Sharath Adavanne , Archontis Politis , Joonas Nikunen , Tuomas Virtanen

The steered response power (SRP) is a popular approach to compute a map of the acoustic scene, typically used for acoustic source localization. The SRP map is obtained as the frequency-weighted output power of a beamformer steered towards a…

音频与语音处理 · 电气工程与系统科学 2024-11-25 Thomas Dietzen , Enzo De Sena , Toon van Waterschoot

This paper addresses the problem of sound-source localization (SSL) with a robot head, which remains a challenge in real-world environments. In particular we are interested in locating speech sources, as they are of high interest for…

声音 · 计算机科学 2020-12-08 Xiaofei Li , Laurent Girin , Fabien Badeig , Radu Horaud

Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global audio and one-scale…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Shentong Mo , Haofan Wang

The task of Visual Sound Source Localization (VSSL) involves identifying the location of sound sources in visual scenes, integrating audio-visual data for enhanced scene understanding. Despite advancements in state-of-the-art (SOTA) models,…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Xavier Juanola , Gloria Haro , Magdalena Fuentes

A method of optimizing secondary source placement in sound field synthesis is proposed. Such an optimization method will be useful when the allowable placement region and available number of loudspeakers are limited. We formulate a…

声音 · 计算机科学 2021-12-14 Keisuke Kimura , Shoichi Koyama , Natsuki Ueno , Hiroshi Saruwatari

This paper addresses the problem of multiple-speaker localization in noisy and reverberant environments, using binaural recordings of an acoustic scene. A Gaussian mixture model (GMM) is adopted, whose components correspond to all the…

声音 · 计算机科学 2017-10-06 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

Steered Response Power (SRP) is a widely used method for the task of sound source localization using microphone arrays, showing satisfactory localization performance on many practical scenarios. However, its performance is diminished under…

声音 · 计算机科学 2024-03-15 Eric Grinstein , Toon van Waterschoot , Mike Brookes , Patrick A. Naylor

Recent Transformer-based 3D object detectors learn point cloud features either from point- or voxel-based representations. However, the former requires time-consuming sampling while the latter introduces quantization errors. In this paper,…

计算机视觉与模式识别 · 计算机科学 2023-05-12 Honghui Yang , Wenxiao Wang , Minghao Chen , Binbin Lin , Tong He , Hua Chen , Xiaofei He , Wanli Ouyang

Localizing linearly moving sound sources using microphone arrays is challenging as the transient nature of the signal leads to relatively short observation periods. Commonly, a moving focus is used and most methods operate at least…

音频与语音处理 · 电气工程与系统科学 2024-08-21 Christian H. Kasess , Wolfgang Kreuzer , Prateek Soni , Holger Waubke

Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two microphone channels. Though DP-RTF fully encodes the sound spatial cues and serves as a reliable localization…

声音 · 计算机科学 2022-02-17 Bing Yang , Hong Liu , Xiaofei Li

We consider the high-resolution imaging problem of 3D point source image recovery from 2D data using a method based on point spread function (PSF) engineering. The method involves a new technique, recently proposed by S.~Prasad, based on…

信号处理 · 电气工程与系统科学 2019-06-13 Chao Wang , Raymond Chan , Mila Nikolova , Robert Plemmons , Sudhakar Prasad

It is commonly believed that multipath hurts various audio processing algorithms. At odds with this belief, we show that multipath in fact helps sound source separation, even with very simple propagation models. Unlike most existing…

声音 · 计算机科学 2019-05-08 Robin Scheibler , Diego Di Carlo , Antoine Deleforge , Ivan Dokmanić

Signal decomposition (SD) approaches aim to decompose non-stationary signals into their constituent amplitude- and frequency-modulated components. This represents an important preprocessing step in many practical signal processing…

信号处理 · 电气工程与系统科学 2022-09-05 Thomas Eriksen , Naveed ur Rehman

While transfer learning is an effective strategy, it often overlooks the opportunity to leverage knowledge from numerous available models online. Addressing this multi-source transfer learning problem is a promising path to boost…

机器学习 · 计算机科学 2026-04-24 Marcin Osial , Bartosz Wójcik , Bartosz Zieliński , Sebastian Cygert

Accurate and early diagnosis of pneumonia through X-ray imaging is essential for effective treatment and improved patient outcomes. Recent advancements in machine learning have enabled automated diagnostic tools that assist radiologists in…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Mete Erdogan , Sebnem Demirtas