English
Related papers

Related papers: Relative Transfer Function Estimation Exploiting S…

200 papers

One of the biggest challenges in multi-microphone applications is the estimation of the parameters of the signal model such as the power spectral densities (PSDs) of the sources, the early (relative) acoustic transfer functions of the…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-16 Andreas I. Koutrouvelis , Richard C. Hendriks , Richard Heusdens , Jesper Jensen

Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Yu Qi , Lipeng Gu , Honghua Chen , Liangliang Nan , Mingqiang Wei

Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for audio source separation, and we call it speaker representation…

Sound · Computer Science 2020-02-28 Seongkyu Mun , Soyeon Choe , Jaesung Huh , Joon Son Chung

Real-world applications such as magnetic resonance imaging with multiple coils, multi-user communication, and diffuse optical tomography often assume a linear model where several sparse signals sharing common sparse supports are acquired by…

Information Theory · Computer Science 2018-10-17 Junan Zhu , Dror Baron

A primary challenge in developing synthetic spatial hearing systems, particularly underwater, is accurately modeling sound scattering. Biological organisms achieve 3D spatial hearing by exploiting sound scattering off their bodies to…

Sound · Computer Science 2026-03-03 Siminfar Samakoush Galougah , Pranav Pulijala , Ramani Duraiswami

This paper addresses the problem of multiple-speaker localization in noisy and reverberant environments, using binaural recordings of an acoustic scene. A Gaussian mixture model (GMM) is adopted, whose components correspond to all the…

Sound · Computer Science 2017-10-06 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

Although the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech is still high. In this paper, we propose a spatio-temporal…

Sound · Computer Science 2021-04-06 Yong Xu , Zhuohuang Zhang , Meng Yu , Shi-Xiong Zhang , Dong Yu

Target-speaker speech recognition aims to recognize target-speaker speech from noisy environments with background noise and interfering speakers. This work presents a joint framework that combines time-domain target-speaker speech…

Sound · Computer Science 2021-03-01 Jiatong Shi , Chunlei Zhang , Chao Weng , Shinji Watanabe , Meng Yu , Dong Yu

End-to-end model, especially Recurrent Neural Network Transducer (RNN-T), has achieved great success in speech recognition. However, transducer requires a great memory footprint and computing time when processing a long decoding sequence.…

Sound · Computer Science 2023-07-18 Xiaohui Zhang , Mangui Liang , Zhengkun Tian , Jiangyan Yi , Jianhua Tao

We compare using a PHOIBLE-based phone mapping method and using phonological features input in transfer learning for TTS in low-resource languages. We use diverse source languages (English, Finnish, Hindi, Japanese, and Russian) and target…

Computation and Language · Computer Science 2023-06-22 Phat Do , Matt Coler , Jelske Dijkstra , Esther Klabbers

Multitrack music transcription aims to transcribe a music audio input into the musical notes of multiple instruments simultaneously. It is a very challenging task that typically requires a more complex model to achieve satisfactory result.…

Sound · Computer Science 2023-06-21 Wei-Tsung Lu , Ju-Chiang Wang , Yun-Ning Hung

In sound field control applications, it is commonly assumed that one has access to an accurate representation of the sound field in the region of interest. This is a problematic assumption since the reconstruction of a sound field from…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 David Sundström , Filip Tronarp , Johan Lindström , Andreas Jakobsson

The recent resurgence of interest in spatio-temporal neural network as speech recognition tool motivates the present investigation. In this paper an approach was developed based on temporal radial basis function "TRBF" looking to many…

Computation and Language · Computer Science 2009-12-22 Mustapha Guezouri , Larbi Mesbahi , Abdelkader Benyettou

Speaker verification systems have seen significant advancements with the introduction of Multi-scale Feature Aggregation (MFA) architectures, such as MFA-Conformer and ECAPA-TDNN. These models leverage information from various network…

Sound · Computer Science 2024-10-08 Satvik Dixit , Massa Baali , Rita Singh , Bhiksha Raj

This paper introduces a novel method to separate noisy speech into low or high frequency frames, in order to improve fundamental frequency (F0) estimation accuracy. In this proposal, the target signal is analyzed by means of the ensemble…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-21 A. Queiroz , R. Coelho

Air pollution, especially particulate matter 2.5 (PM2.5), is a pressing concern for public health and is difficult to estimate in developing countries (data-poor regions) due to a lack of ground sensors. Transfer learning models can be…

Machine Learning · Computer Science 2024-06-25 Shrey Gupta , Yongbee Park , Jianzhao Bi , Suyash Gupta , Andreas Züfle , Avani Wildani , Yang Liu

Many super-resolution (SR) models are optimized for high performance only and therefore lack efficiency due to large model complexity. As large models are often not practical in real-world applications, we investigate and propose novel loss…

Image and Video Processing · Electrical Eng. & Systems 2021-06-03 Dario Fuoli , Luc Van Gool , Radu Timofte

Accurate channel estimation is critical to fully exploit the beamforming gains when communicating with extremely large aperture arrays. The propagation distances between the user and receiver, which potentially has thousands of…

Signal Processing · Electrical Eng. & Systems 2024-02-01 Alva Kosasih , Özlem Tuğfe Demir , Emil Björnson

Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal enhancement. However, the textbook solutions for optimal…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Reinhold Haeb-Umbach , Tomohiro Nakatani , Marc Delcroix , Christoph Boeddeker , Tsubasa Ochiai

Scene flow estimation is an essential ingredient for a variety of real-world applications, especially for autonomous agents, such as self-driving cars and robots. While recent scene flow estimation approaches achieve a reasonable accuracy,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-07 Yushan Zhang , Bastian Wandt , Maria Magnusson , Michael Felsberg
‹ Prev 1 8 9 10 Next ›