English
Related papers

Related papers: Binaural Source Localization based on Modulation-D…

200 papers

In this thesis, we propose an artificial auditory system that gives a robot the ability to locate and track sounds, as well as to separate simultaneous sound sources and recognising simultaneous speech. We demonstrate that it is possible to…

Robotics · Computer Science 2016-02-23 Jean-Marc Valin

This paper presents SSLIDE, Sound Source Localization for Indoors using DEep learning, which applies deep neural networks (DNNs) with encoder-decoder structure to localize sound sources with random positions in a continuous space. The…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-17 Yifan Wu , Roshan Ayyalasomayajula , Michael J. Bianco , Dinesh Bharadia , Peter Gerstoft

Recent advances in multi-modal large language models (MLLMs) have opened new possibilities for unified modeling of speech, text, images, and other modalities. Building on our prior work, this paper examines the conditions and model…

Sound · Computer Science 2025-07-28 Yiwen Guan , Viet Anh Trinh , Vivek Voleti , Jacob Whitehill

We consider a set of M images, whose pixel intensities at a common point can be treated as the components of a M-dimensional vector. We are interested in the estimation of the modulus of such a vector associated to a compact source. For…

Cosmology and Nongalactic Astrophysics · Physics 2011-01-05 F. Argueso , J. L. Sanz , D. Herranz

In this paper, we propose a deep learning based multi-speaker direction of arrival (DOA) estimation with audio and visual signals by using permutation-free loss function. We first collect a data set for multi-modal sound source localization…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-27 Qing Wang , Hang Chen , Ya Jiang , Zhe Wang , Yuyang Wang , Jun Du , Chin-Hui Lee

Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize…

Sound · Computer Science 2025-08-07 Yu Chen , Hongxu Zhu , Jiadong Wang , Kainan Chen , Xinyuan Qian

Audio-visual segmentation (AVS) is a challenging task that involves accurately segmenting sounding objects based on audio-visual cues. The effectiveness of audio-visual learning critically depends on achieving accurate cross-modal alignment…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Yuanhong Chen , Yuyuan Liu , Hu Wang , Fengbei Liu , Chong Wang , Helen Frazer , Gustavo Carneiro

High spatial frequency information, including fine details like textures, significantly contributes to the accuracy of semantic segmentation. However, according to the Nyquist-Shannon Sampling Theorem, high-frequency components are…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Linwei Chen , Ying Fu , Lin Gu , Dezhi Zheng , Jifeng Dai

Existing machine learning research has achieved promising results in monaural audio-visual separation (MAVS). However, most MAVS methods purely consider what the sound source is, not where it is located. This can be a problem in VR/AR…

Sound · Computer Science 2023-11-01 Yuxin Ye , Wenming Yang , Yapeng Tian

It is known that adverse environments such as high reverberation and low signal-to-noise ratio (SNR) pose a great challenge to indoor sound source localization. To address this challenge, in this paper, we propose a sound source…

Sound · Computer Science 2018-12-05 Yingxiang Sun , Jiajia Chen , Chau Yuen , Susanto Rahardja

We consider the localization problem of multiple wideband sources in a multi-path environment by coherently taking into account the attenuation characteristics and the time delays in the reception of the signal. Our proposed method leaves…

Networking and Internet Architecture · Computer Science 2015-05-27 Hamidreza Aghasi , Hamidreza Amindavar , Alireza Aghasi

To enhance immersive experiences, binaural audio offers spatial awareness of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural audio to binaural…

Sound · Computer Science 2025-06-03 Tianrui Pan , Jie Liu , Zewen Huang , Jie Tang , Gangshan Wu

We propose a novel multi-source direction of arrival (DOA) estimation technique using a convolutional neural network algorithm which learns the modal coherence patterns of an incident soundfield through measured spherical harmonic…

Sound · Computer Science 2020-03-19 A. Fahim , P. N. Samarasinghe , T. D. Abhayapala

In this paper, the problem of determining the number of signal sources impinging on an array of sensors and estimating their directions-of-arrival (DOAs) in the presence of spatially white nonuniform noise is considered. It is known that,…

Signal Processing · Electrical Eng. & Systems 2021-09-10 Mahmood Karimi

An anomalous sound detection system to detect unknown anomalous sounds usually needs to be built using only normal sound data. Moreover, it is desirable to improve the system by effectively using a small amount of anomalous sound data,…

Sound · Computer Science 2021-06-14 Ibuki Kuroyanagi , Tomoki Hayashi , Kazuya Takeda , Tomoki Toda

We consider the acoustic source imaging problems using multiple frequency data. Using the data of one observation direction/point, we prove that some information (size and location) of the source support can be recovered. A non-iterative…

Analysis of PDEs · Mathematics 2017-12-08 Ala Alzaalig , Guanghui Hu , Xiaodong Liu , Jiguang Sun

This paper is about alerting acoustic event detection and sound source localisation in an urban scenario. Specifically, we are interested in spotting the presence of horns, and sirens of emergency vehicles. In order to obtain a reliable…

Sound · Computer Science 2022-03-29 Letizia Marchegiani , Paul Newman

As speech processing systems in mobile and edge devices become more commonplace, the demand for unintrusive speech quality monitoring increases. Deep learning methods provide high-quality estimates of objective and subjective speech quality…

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Qiang Huang , Thomas Hain

This paper investigates the transmit beamforming design for multiple-input multiple-output systems to support both multi-target localization and multi-user communications. To enhance the target localization performance, we derive the…

Signal Processing · Electrical Eng. & Systems 2025-01-14 Bo Tang , Da Li , Wenjun Wu , Astha Saini , Prabhu Babu , Petre Stoica