English
Related papers

Related papers: Latent Acoustic Mapping for Direction of Arrival E…

200 papers

In this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked…

Sound · Computer Science 2024-10-18 Ashish Seth , Ramaneswaran Selvakumar , S Sakshi , Sonal Kumar , Sreyan Ghosh , Dinesh Manocha

Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g.,…

Sound · Computer Science 2025-12-15 Hualei Wang , Yiming Li , Shuo Ma , Hong Liu , Xiangdong Wang

Unmanned Aerial Vehicles (UAVs) are increasingly deployed in search-and-rescue (SAR) missions, yet continuous and reliable victim detection and localization remain challenging due to on-board hardware constraints. This paper designs an…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Yi Hong , Mingyang Wang , Yalin Liu , Yaru Fu , Kevin Hung

Deep learning-based sound event localization and classification is an emerging research area within wireless acoustic sensor networks. However, current methods for sound event localization and classification typically rely on a single…

Deep neural networks have recently led to promising results for the task of multiple sound source localization. Yet, they require a lot of training data to cover a variety of acoustic conditions and microphone array layouts. One can…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-18 Guillaume Le Moing , Phongtharin Vinayavekhin , Don Joven Agravante , Tadanobu Inoue , Jayakorn Vongkulbhisal , Asim Munawar , Ryuki Tachibana

In this paper, we propose an artificial intelligence (AI)-enhanced hybrid simultaneous localization and mapping (SLAM) method that performs Bayesian inference directly on raw radio-frequency (RF) signals while learning an environment model…

Signal Processing · Electrical Eng. & Systems 2026-03-17 Alexander Venus , Benjamin Deutschmann , Alexander Fuchs , Christian Knoll , Erik Leitinger

Most existing masked audio modeling (MAM) methods learn audio representations by masking and reconstructing local spectrogram patches. However, the reconstruction loss mainly accounts for the signal-level quality of the reconstructed…

Sound · Computer Science 2024-01-30 Yifei Xin , Xiulian Peng , Yan Lu

Machine learning and especially deep learning have garneredtremendous popularity in recent years due to their increased performanceover other methods. The availability of large amount of data has aidedin the progress of deep learning.…

Machine Learning · Computer Science 2019-09-06 Sharath M. Shankaranarayana , Davor Runje

Passive Acoustic Monitoring (PAM) analysis is often hindered by the intensive manual effort needed to create labelled training data. This study introduces a synthetic data framework to generate large volumes of richly labelled training data…

Sound · Computer Science 2025-07-23 Kaspar Soltero , Tadeu Siqueira , Stefanie Gutschmidt

Integrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial…

Sound · Computer Science 2025-04-28 Ayushi Mishra , Yang Bai , Priyadarshan Narayanasamy , Nakul Garg , Nirupam Roy

Deep neural networks (DNNs) typically employ an end-to-end (E2E) training paradigm which presents several challenges, including high GPU memory consumption, inefficiency, and difficulties in model parallelization during training. Recent…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Yuming Zhang , Shouxin Zhang , Peizhe Wang , Feiyu Zhu , Dongzhi Guan , Junhao Su , Jiabin Liu , Changpeng Cai

Passive Acoustic Monitoring (PAM) is an efficient and non-invasive method for surveying ecosystems at a reduced cost. Typically, autonomous recorders allow the acquisition of vast bioacoustic datasets which are then analyzed. However, power…

Sound · Computer Science 2026-05-06 Louis Lerbourg , Paul Peyret , Juliette Linossier , Marielle Malfante

Models of acoustic word embeddings (AWEs) learn to map variable-length spoken word segments onto fixed-dimensionality vector representations such that different acoustic exemplars of the same word are projected nearby in the embedding…

Computation and Language · Computer Science 2022-09-20 Badr M. Abdullah , Bernd Möbius , Dietrich Klakow

Simultaneous localization and mapping (SLAM) is the process of constructing a global model of an environment from local observations of it; this is a foundational capability for mobile robots, supporting such core functions as planning,…

Robotics · Computer Science 2021-03-10 David M. Rosen , Kevin J. Doherty , Antonio Teran Espinoza , John J. Leonard

Long-context understanding is crucial for many NLP applications, yet transformers struggle with efficiency due to the quadratic complexity of self-attention. Sparse attention methods alleviate this cost but often impose static, predefined…

Computation and Language · Computer Science 2025-06-16 Hanzhi Zhang , Heng Fan , Kewei Sha , Yan Huang , Yunhe Feng

Localizing partial deepfake audio, where only segments of speech are manipulated, remains challenging due to the subtle and scattered nature of these modifications. Existing approaches typically rely on frame-level predictions to identify…

Sound · Computer Science 2026-01-30 Yuchen Mao , Wen Huang , Yanmin Qian

Understanding the internal representations of large language models (LLMs) remains a central challenge for interpretability research. Sparse autoencoders (SAEs) offer a promising solution by decomposing activations into interpretable…

Machine Learning · Computer Science 2025-10-10 Yifei Yao , Mengnan Du

In this paper, we propose a deep learning based multi-speaker direction of arrival (DOA) estimation with audio and visual signals by using permutation-free loss function. We first collect a data set for multi-modal sound source localization…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-27 Qing Wang , Hang Chen , Ya Jiang , Zhe Wang , Yuyang Wang , Jun Du , Chin-Hui Lee

Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Published approaches that consider both audio and words/tags…

Sound · Computer Science 2020-10-28 Xavier Favory , Konstantinos Drossos , Tuomas Virtanen , Xavier Serra