中文
相关论文

相关论文: HRTF measurement for accurate sound localization c…

200 篇论文

Sound-tracking refers to the process of determining the direction from which a sound originates, making it a fundamental component of sound source localization. This capability is essential in a variety of applications, including security…

声音 · 计算机科学 2025-10-13 Mahdi Ali Pour , Zahra Habibzadeh

Localizing linearly moving sound sources using microphone arrays is challenging as the transient nature of the signal leads to relatively short observation periods. Commonly, a moving focus is used and most methods operate at least…

音频与语音处理 · 电气工程与系统科学 2024-08-21 Christian H. Kasess , Wolfgang Kreuzer , Prateek Soni , Holger Waubke

Binaural reproduction aims to deliver immersive spatial audio with high perceptual realism over headphones. Loss functions play a central role in optimizing and evaluating algorithms that generate binaural signals. However, traditional…

音频与语音处理 · 电气工程与系统科学 2026-04-03 Boaz Rafaely , Stefan Weinzierl , Or Berebi , Fabian Brinkmann

Learning to localize the sound source in videos without explicit annotations is a novel area of audio-visual research. Existing work in this area focuses on creating attention maps to capture the correlation between the two modalities to…

计算机视觉与模式识别 · 计算机科学 2022-11-08 Dennis Fedorishin , Deen Dayal Mohan , Bhavin Jawade , Srirangaraj Setlur , Venu Govindaraju

In this work, we introduce HeFT (Head-Frequency Tracker), a zero-shot point tracking framework that leverages the visual priors of pretrained video diffusion models. To better understand how they encode spatiotemporal information, we…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Tianyu Yuan , Yuanbo Yang , Lin-Zhuo Chen , Yao Yao , Zhuzhong Qian

Automatic speech recognition (ASR) on multi-talker recordings is challenging. Current methods using 3D spatial data from multi-channel audio and visual cues focus mainly on direct waves from the target speaker, overlooking reflection wave…

音频与语音处理 · 电气工程与系统科学 2024-06-13 Yiwen Shao , Shi-Xiong Zhang , Dong Yu

This paper presents Rec-RIR for monaural blind room impulse response (RIR) identification. Rec-RIR is developed based on the convolutive transfer function (CTF) approximation, which models reverberation effect within narrow-band filter…

音频与语音处理 · 电气工程与系统科学 2026-01-22 Pengyu Wang , Xiaofei Li

An analysis of the relationship between the bandwidth of acoustic signals and the required resolution of steered-response power phase transform (SRP-PHAT) maps used for sound source localization is presented. This relationship does not rely…

In this paper, we attempt to study the conditioning of the Spherical Harmonic Matrix (SHM), which is widely used in the discrete, limited order orthogonal representation of sound fields. SHM's has been widely used in the audio applications…

音频与语音处理 · 电气工程与系统科学 2018-03-07 C Sandeep Reddy , Rajesh M Hegde

This paper addresses the problem of infants' cry fundamental frequency estimation. The fundamental frequency is estimated using a modified simple inverse filtering tracking (SIFT) algorithm. The performance of the modified SIFT is studied…

声音 · 计算机科学 2010-09-16 Dror Lederman

Acoustic signal processing in the spherical harmonics domain (SHD) is an active research area that exploits the signals acquired by higher order microphone arrays. A very important task is that concerning the localization of active sound…

音频与语音处理 · 电气工程与系统科学 2024-01-29 Maximo Cobos , Mirco Pezzoli , Fabio Antonacci , Augusto Sarti

Measuring room impulse responses (RIRs) at multiple spatial points is a time-consuming task, while simulations require detailed knowledge of the room's acoustic environment. In prior work, we proposed a method for estimating the early part…

音频与语音处理 · 电气工程与系统科学 2025-06-16 Kathleen MacWilliam , Thomas Dietzen , Toon van Waterschoot

Determining the head orientation of a talker is not only beneficial for various speech signal processing applications, such as source localization or speech enhancement, but also facilitates intuitive voice control and interaction with…

音频与语音处理 · 电气工程与系统科学 2026-02-10 Kaspar Müller , Bilgesu Çakmak , Paul Didier , Simon Doclo , Jan Østergaard , Tobias Wolff

During the last years the transfer of frequency signals through optical fibers has shown ultra low instabilities in various configurations. The outstanding experimental results of such point-to-point connections is motivation to develop a…

光学 · 物理学 2019-10-25 M. Rost , M. Fujieda , D. Piester

Recent advances in generative speech have increased the need for automatic detection of obviously failed synthetic outputs. This is particularly important in clinical settings such as AVATAR therapy, in which schizophrenia patients engage…

音频与语音处理 · 电气工程与系统科学 2026-05-12 Jana Shokr , Minos Papadopoulos , Jeremy Cooperstock , Pavo Orepic

Simultaneous space-time focusing (SSTF) is sometimes claimed to reduce the longitudinal extent of the high-intensity region near the focus, in contradiction to the original work on this topic. Here we seek to address this confusion by using…

光学 · 物理学 2025-02-19 Emily Archer , Bangshan Sun , Roman Walczak , Martin Booth , Simon Hooker

Multi-channel speech enhancement utilizes spatial information from multiple microphones to extract the target speech. However, most existing methods do not explicitly model spatial cues, instead relying on implicit learning from…

声音 · 计算机科学 2023-09-20 Jiahui Pan , Shulin He , Hui Zhang , Xueliang Zhang

Owning to the reflection gain and double path loss featured by intelligent reflecting surface (IRS) channels, handover (HO) locations become irregular and the signal strength fluctuates sharply with variations in IRS connections during HO,…

信号处理 · 电气工程与系统科学 2025-01-28 Haoyan Wei , Hongtao Zhang

To estimate the direction of arrival (DOA) of multiple speakers with methods that use prototype transfer functions, frequency-dependent spatial spectra (SPS) are usually constructed. To make the DOA estimation robust, SPS from different…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Daniel Fejgin , Elior Hadad , Sharon Gannot , Zbyněk Koldovský , Simon Doclo

This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…

声音 · 计算机科学 2019-01-31 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud