English
Related papers

Related papers: Bias Analysis of Spatial Coherence-Based RTF Vecto…

200 papers

In indoor scenes, reverberation is a crucial factor in degrading the perceived quality and intelligibility of speech. In this work, we propose a generative dereverberation method. Our approach is based on a probabilistic model utilizing a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-18 Pengyu Wang , Xiaofei Li

Personalized head-related transfer functions (HRTFs) are essential for ensuring a realistic auditory experience over headphones, because they take into account individual anatomical differences that affect listening. Most machine learning…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 You Zhang , Andrew Francl , Ruohan Gao , Paul Calamia , Zhiyao Duan , Ishwarya Ananthabhotla

Existing speaker verification (SV) systems often suffer from performance degradation if there is any language mismatch between model training, speaker enrollment, and test. A major cause of this degradation is that most existing SV methods…

Sound · Computer Science 2017-06-27 Lantian Li , Dong Wang , Askar Rozi , Thomas Fang Zheng

Virtual sensing (VS) technology enables active noise control (ANC) systems to attenuate noise at virtual locations distant from the physical error microphones. Appropriate auxiliary filters (AF) can significantly enhance the effectiveness…

Signal Processing · Electrical Eng. & Systems 2024-09-10 Boxiang Wang , Dongyuan Shi , Zhengding Luo , Xiaoyi Shen , Junwei Ji , Woon-Seng Gan

This paper proposes a neural network based speech separation method using spatially distributed microphones. Unlike with traditional microphone array settings, neither the number of microphones nor their spatial arrangement is known in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Dongmei Wang , Zhuo Chen , Takuya Yoshioka

Supervised learning based methods for source localization, being data driven, can be adapted to different acoustic conditions via training and have been shown to be robust to adverse acoustic environments. In this paper, a convolutional…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-22 Soumitro Chakrabarty , Emanuël A. P. Habets

Automatic speaker verification (ASV) technology is recently finding its way to end-user applications for secure access to personal data, smart services or physical facilities. Similar to other biometric technologies, speaker verification is…

Sound · Computer Science 2016-09-16 Cemal Hanilci , Tomi Kinnunen , Md Sahidullah , Aleksandr Sizov

In this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T based SCD model loss optimizes all output tokens equally. Due to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-06 Guanlong Zhao , Quan Wang , Han Lu , Yiling Huang , Ignacio Lopez Moreno

Despite growing attention to deepfake speech detection, the aspects of bias and fairness remain underexplored in the speech domain. To address this gap, we introduce the Speaker Characteristics Deepfake (SCDF) dataset: a novel, richly…

Sound · Computer Science 2025-08-12 Vojtěch Staněk , Karel Srna , Anton Firc , Kamil Malinka

Self-Supervised Learning (SSL) Automatic Speech Recognition (ASR) models have shown great promise over Supervised Learning (SL) ones in low-resource settings. However, the advantages of SSL are gradually weakened when the amount of labeled…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 Li Fu , Siqi Li , Qingtao Li , Fangzhu Li , Liping Deng , Lu Fan , Meng Chen , Youzheng Wu , Xiaodong He

The second-order harmonic (2f) component generated by twin-rotary compressor is a dominant low-frequency noise source of variable refrigerant flow (VRF) outdoor units, yet its amplitude fluctuates strongly with environmental thermal load…

Signal Processing · Electrical Eng. & Systems 2026-05-05 ZhiWei Su , Ding Wang , Yuan Guo , Yang Qiao , HongJun Cao

The paper presents a method for audio-based vehicle counting (VC) in low-to-moderate traffic using one-channel sound. We formulate VC as a regression problem, i.e., we predict the distance between a vehicle and the microphone. Minima of the…

Sound · Computer Science 2020-10-23 Slobodan Djukanović , Jiři Matas , Tuomas Virtanen

We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed for the generative…

Sound · Computer Science 2023-10-27 Ali Golmakani , Mostafa Sadeghi , Xavier Alameda-Pineda , Romain Serizel

Multi-channel short-time Fourier transform (STFT) domain-based processing of reverberant microphone signals commonly relies on power-spectral-density (PSD) estimates of early source images, where early refers to reflections contained within…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-21 T. Dietzen , S. Doclo , M. Moonen , T. van Waterschoot

This paper proposes singing voice synthesis (SVS) based on frame-level sequence-to-sequence models considering vocal timing deviation. In SVS, it is essential to synchronize the timing of singing with temporal structures represented by…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-23 Miku Nishihara , Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda

Sound field estimation with moving microphones can increase flexibility, decrease measurement time, and reduce equipment constraints compared to using stationary microphones. In this paper a sound field estimation method based on kernel…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-23 Jesper Brunnström , Martin Bo Møller , Jan Østergaard , Shoichi Koyama , Toon van Waterschoot , Marc Moonen

Sentiment Analysis is a branch of Affective Computing usually considered a binary classification task. In this line of reasoning, Sentiment Analysis can be applied in several contexts to classify the attitude expressed in text samples, for…

Information Retrieval · Computer Science 2020-08-13 Flavio Carvalho , Gustavo Paiva Guedes

Multiple moving sound source localization in real-world scenarios remains a challenging issue due to interaction between sources, time-varying trajectories, distorted spatial cues, etc. In this work, we propose to use deep learning…

Sound · Computer Science 2022-02-17 Bing Yang , Hong Liu , Xiaofei Li

The acquisition of high-quality labeled synthetic aperture radar (SAR) data is challenging due to the demanding requirement for expert knowledge. Consequently, the presence of unreliable noisy labels is unavoidable, which results in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Yimin Fu , Zhunga Liu , Dongxiu Guo , Longfei Wang

A method of interpolating the acoustic transfer function (ATF) between regions that takes into account both the physical properties of the ATF and the directionality of region configurations is proposed. Most spatial ATF interpolation…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-06 Juliano G. C. Ribeiro , Shoichi Koyama , Hiroshi Saruwatari