English
Related papers

Related papers: HRTF measurement for accurate sound localization c…

200 papers

Intelligent spectrum management is crucial for improving spectrum efficiency and achieving secure utilization of spectrum resources. However, existing intelligent spectrum management methods, typically based on small-scale models, suffer…

Signal Processing · Electrical Eng. & Systems 2025-12-16 Fuhui Zhou , Chunyu Liu , Hao Zhang , Wei Wu , Qihui Wu , Tony Q. S. Quek , Chan-Byoung Chae

Target source extraction is significant for improving human speech intelligibility and the speech recognition performance of computers. This study describes a method for target source extraction, called the similarity-and-independence-aware…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-22 Atsuo Hiroe

An interpolation method for region-to-region acoustic transfer functions (ATFs) based on kernel ridge regression with an adaptive kernel is proposed. Most current ATF interpolation methods do not incorporate the acoustic properties for…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-08 Juliano G. C. Ribeiro , Shoichi Koyama , Hiroshi Saruwatari

Neural HMMs are a type of neural transducer recently proposed for sequence-to-sequence modelling in text-to-speech. They combine the best features of classic statistical speech synthesis and modern neural TTS, requiring less data and fewer…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Shivam Mehta , Ambika Kirkland , Harm Lameris , Jonas Beskow , Éva Székely , Gustav Eje Henter

Conventional sound source localization methods are mostly based on a single microphone array that consists of multiple microphones. They are usually formulated as the estimation of the direction of arrival problem. In this paper, we propose…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-18 Yijun Gong , Shupei Liu , Xiao-Lei Zhang

Speech translation (ST) automatically converts utterances in a source language into text in another language. Splitting continuous speech into shorter segments, known as speech segmentation, plays an important role in ST. Recent…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-19 Ryo Fukuda , Katsuhito Sudoh , Satoshi Nakamura

We introduce HiFi-HARP, a large-scale dataset of 7th-order Higher-Order Ambisonic Room Impulse Responses (HOA-RIRs) consisting of more than 100,000 RIRs generated via a hybrid acoustic simulation in realistic indoor scenes. HiFi-HARP…

Sound · Computer Science 2025-10-27 Shivam Saini , Jürgen Peissig

The introduction of ISO 12913-2:2018 has provided a framework for standardized data collection and reporting procedures for soundscape practitioners. A strong emphasis was placed on the use of calibrated head and torso simulators (HATS) for…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-11 Bhan Lam , Kenneth Ooi , Karn N. Watcharasupat , Zhen-Ting Ong , Yun-Ting Lau , Trevor Wong , Woon-Seng Gan

Estimation of the location of sound sources is usually done using microphone arrays. Such settings provide an environment where we know the difference between the received signals among different microphones in the terms of phase or…

Sound · Computer Science 2026-01-08 Helena Peic Tukuljac , Herve Lissek , Pierre Vandergheynst

Heart sound auscultation holds significant importance in the diagnosis of congenital heart disease. However, existing methods for Heart Sound Diagnosis (HSD) tasks are predominantly limited to a few fixed categories, framing the HSD task as…

Sound · Computer Science 2024-08-19 Zihan Zhao , Pingjie Wang , Liudan Zhao , Yuchen Yang , Ya Zhang , Kun Sun , Xin Sun , Xin Zhou , Yu Wang , Yanfeng Wang

Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embedding vector each,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Inho Kim , Youngkil Song , Jicheol Park , Won Hwa Kim , Suha Kwak

Hearing aids are typically equipped with multiple microphones to exploit spatial information for source localisation and speech enhancement. Especially for hearing aids, a good source localisation is important: it not only guides source…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 Siyuan Song , Stijn Kindt , Jasper Maes , Alexander Bohlender. Nilesh Madhu

A new impulse response (IR) dataset called "MeshRIR" is introduced. Currently available datasets usually include IRs at an array of microphones from several source positions under various room conditions, which are basically designed for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-26 Shoichi Koyama , Tomoya Nishida , Keisuke Kimura , Takumi Abe , Natsuki Ueno , Jesper Brunnström

Normal-hearing listeners adapt to alterations in sound localization cues. This adaptation can result from the establishment of a new spatial map of the altered cues or from a stronger relative weighting of unaltered compared to altered…

Neurons and Cognition · Quantitative Biology 2021-05-10 Maike Klingel , Norbert Kopco , Bernhard Laback

Many spatial filtering algorithms used for voice capture in, e.g., teleconferencing applications, can benefit from or even rely on knowledge of Relative Transfer Functions (RTFs). Accordingly, many RTF estimators have been proposed which,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-06 Andreas Brendel , Johannes Zeitler , Walter Kellermann

Sound source localization (SSL) technology plays a crucial role in various application areas such as fault diagnosis, speech separation, and vibration noise reduction. Although beamforming algorithms are widely used in SSL, their resolution…

Sound · Computer Science 2024-10-01 Wenbo Ma , Yan Lu , Yijun Liu

Room impulse response (RIR), which measures the sound propagation within an environment, is critical for synthesizing high-fidelity audio for a given environment. Some prior work has proposed representing RIR as a neural field function of…

Sound · Computer Science 2023-09-29 Susan Liang , Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu

The audio spectrogram is a time-frequency representation that has been widely used for audio classification. One of the key attributes of the audio spectrogram is the temporal resolution, which depends on the hop size used in the Short-Time…

Sound · Computer Science 2024-01-15 Haohe Liu , Xubo Liu , Qiuqiang Kong , Wenwu Wang , Mark D. Plumbley

Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Rahil Parikh , Nadee Seneviratne , Ganesh Sivaraman , Shihab Shamma , Carol Espy-Wilson

Source-free unsupervised domain adaptation (SFUDA) aims to learn a target domain model using unlabeled target data and the knowledge of a well-trained source domain model. Most previous SFUDA works focus on inferring semantics of target…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Jiangbo Pei , Zhuqing Jiang , Aidong Men , Liang Chen , Yang Liu , Qingchao Chen