English
Related papers

Related papers: Gaussian Process Models for HRTF based Sound-Sourc…

200 papers

Audio deepfake model attribution aims to mitigate the misuse of synthetic speech by identifying the source model responsible for generating a given audio sample, enabling accountability and informing vendors. The task is challenging, but…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Gabriel Pîrlogeanu , Adriana Stan , Horia Cucu

Wearable sensors enable the continuous acquisition of high-resolution physiological waveforms, such as photoplethysmography and accelerometry, under free-living conditions. However, inferring health-related phenotypes from these signals…

Virtual sound synthesis is a technology that allows users to perceive spatial sound through headphones or earphones. However, accurate virtual sound requires an individual head-related transfer function (HRTF), which can be difficult to…

Sound · Computer Science 2023-10-24 Tatsuki Kobayashi , Yoshiko Maruyama , Isao Nambu , Shohei Yano , Yasuhiro Wada

Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living environments, labeling…

Sound · Computer Science 2020-07-29 Yoshiki Masuyama , Yoshiaki Bando , Kohei Yatabe , Yoko Sasaki , Masaki Onishi , Yasuhiro Oikawa

Head-related transfer functions (HRTFs) are crucial for spatial soundfield reproduction in virtual reality applications. However, obtaining personalized, high-resolution HRTFs is a time-consuming and costly task. Recently, deep…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Xingyu Chen , Fei Ma , Yile Zhang , Amy Bastine , Prasanga N. Samarasinghe

In Self-Supervised Learning (SSL), pre-training and evaluation are resource intensive. In the speech domain, current indicators of the quality of SSL models during pre-training, such as the loss, do not correlate well with downstream…

Sound · Computer Science 2025-06-03 Ryan Whetten , Lucas Maison , Titouan Parcollet , Marco Dinarelli , Yannick Estève

Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representations using unlabeled speech for content-related tasks is…

Computation and Language · Computer Science 2024-06-14 Amit Meghanani , Thomas Hain

Deep Learning models have become potential candidates for auditory neuroscience research, thanks to their recent successes on a variety of auditory tasks. Yet, these models often lack interpretability to fully understand the exact…

Sound · Computer Science 2021-08-04 Rachid Riad , Julien Karadayi , Anne-Catherine Bachoud-Lévi , Emmanuel Dupoux

Learning to localize the sound source in videos without explicit annotations is a novel area of audio-visual research. Existing work in this area focuses on creating attention maps to capture the correlation between the two modalities to…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Dennis Fedorishin , Deen Dayal Mohan , Bhavin Jawade , Srirangaraj Setlur , Venu Govindaraju

Automatic singing voice understanding tasks, such as singer identification, singing voice transcription, and singing technique classification, benefit from data-driven approaches that utilize deep learning techniques. These approaches work…

Sound · Computer Science 2023-09-06 Yuya Yamamoto

Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial upsampling has shown remarkable progress with neural fields,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-23 Yoshiki Masuyama , Gordon Wichern , François G. Germain , Christopher Ick , Jonathan Le Roux

The existing source cell-phone recognition method lacks the long-term feature characterization of the source device, resulting in inaccurate representation of the source cell-phone related features which leads to insufficient recognition…

Sound · Computer Science 2022-08-29 Chunyan Zeng , Shixiong Feng , Zhifeng Wang , Xiangkui Wan , Yunfan Chen , Nan Zhao

Speaker localization for binaural microphone arrays has been widely studied for applications such as speech communication, video conferencing, and robot audition. Many methods developed for this task, including the direct path dominance…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-01 Yanir Maymon , Israel Nelken , Boaz Rafaely

To achieve immersive spatial audio rendering on VR/AR devices, high-quality Head-Related Transfer Functions (HRTFs) are essential. In general, HRTFs are subject-dependent and position-dependent, and their measurement is time-consuming and…

Sound · Computer Science 2025-11-17 De Hu , Junsheng Hu , Cuicui Jiang

This paper presents a sound source localization strategy that relies on a microphone array embedded in an unmanned ground vehicle and an asynchronous close-talking microphone near the operator. A signal coarse alignment strategy is combined…

Robotics · Computer Science 2025-07-30 Victor Liu , Timothy Du , Jordy Sehn , Jack Collier , François Grondin

Self-Supervised Learning (SSL) is crucial for real-world applications, especially in data-hungry domains such as healthcare and self-driving cars. In addition to a lack of labeled data, these applications also suffer from distributional…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 Ha Manh Bui , Iliana Maifeld-Carucci

Gaussian process (GP) audio source separation is a time-domain approach that circumvents the inherent phase approximation issue of spectrogram based methods. Furthermore, through its kernel, GPs elegantly incorporate prior knowledge about…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-22 Pablo A. Alvarado , Mauricio A. Álvarez , Dan Stowell

This article is a survey on deep learning methods for single and multiple sound source localization. We are particularly interested in sound source localization in indoor/domestic environment, where reverberation and diffuse noise are…

Sound · Computer Science 2022-07-20 Pierre-Amaury Grumiaux , Srđan Kitić , Laurent Girin , Alexandre Guérin

Expressing head-related transfer functions (HRTFs) in spherical harmonic (SH) domain has been thoroughly studied as a method of obtaining continuity over space. However, HRTFs are functions not only of direction but also of frequency. This…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-13 Adam Szwajcowski

In many multi-microphone algorithms, an estimate of the relative transfer functions (RTFs) of the desired speaker is required. Recently, a computationally efficient RTF vector estimation method was proposed for acoustic sensor networks,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-07 Wiebke Middelberg , Simon Doclo