中文
相关论文

相关论文: Learning Magnitude Distribution of Sound Fields vi…

200 篇论文

This paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used because it works…

Compensation for channel mismatch and noise interference is essential for robust automatic speech recognition. Enhanced speech has been introduced into the multi-condition training of acoustic models to improve their generalization ability.…

声音 · 计算机科学 2022-11-24 Hung-Shin Lee , Pin-Yuan Chen , Yao-Fei Cheng , Yu Tsao , Hsin-Min Wang

The sampling of sound fields involves the measurement of spatially dependent room impulse responses, where the Nyquist-Shannon sampling theorem applies in both the temporal and spatial domain. Therefore, sampling inside a volume of interest…

声音 · 计算机科学 2016-09-30 Fabrice Katzberg , Radoslaw Mazur , Marco Maass , Philipp Koch , Alfred Mertins

Most soundfield synthesis approaches deal with extensive and regular loudspeaker arrays, which are often not suitable for home audio systems, due to physical space constraints. In this article we propose a technique for soundfield synthesis…

音频与语音处理 · 电气工程与系统科学 2024-07-09 Luca Comanducci , Fabio Antonacci , Augusto Sarti

Context propagation remains a central challenge in language model architectures, particularly in tasks requiring the retention of long-range dependencies. Conventional attention mechanisms, while effective in many applications, exhibit…

计算与语言 · 计算机科学 2025-03-26 Alfred Bexley , Lukas Radcliffe , Giles Weatherstone , Joseph Sakau

A method is presented for estimating and reconstructing the sound field within a room using physics-informed neural networks. By incorporating a limited set of experimental room impulse responses as training data, this approach combines…

音频与语音处理 · 电气工程与系统科学 2024-01-03 Xenofon Karakonstantis , Diego Caviedes-Nozal , Antoine Richard , Efren Fernandez-Grande

In most current approaches of speech processing, information is extracted from the magnitude spectrum. However recent perceptual studies have underlined the importance of the phase component. The goal of this paper is to investigate the…

声音 · 计算机科学 2020-01-03 Thomas Drugman , Thomas Dubuisson , Thierry Dutoit

We propose a contrastive conditional latent diffusion model for audio-visual segmentation (AVS) to thoroughly investigate the impact of audio, where the correlation between audio and the final segmentation map is modeled to guarantee the…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Yuxin Mao , Jing Zhang , Mochu Xiang , Yunqiu Lv , Dong Li , Yiran Zhong , Yuchao Dai

Accurately interpreting cardiac auscultation signals plays a crucial role in diagnosing and managing cardiovascular diseases. However, the paucity of labelled data inhibits classification models' training. Researchers have turned to…

声音 · 计算机科学 2025-06-18 Leigh Abbott , Milan Marocchi , Matthew Fynn , Yue Rong , Sven Nordholm

As the capabilities of large-scale pre-trained models evolve, understanding the determinants of their outputs becomes more important. Feature attribution aims to reveal which parts of the input elements contribute the most to model outputs.…

计算与语言 · 计算机科学 2025-05-23 Gaofei Shen , Hosein Mohebbi , Arianna Bisazza , Afra Alishahi , Grzegorz Chrupała

Fr\'echet Audio Distance (FAD) is the de facto standard for evaluating text-to-audio generation, yet its scores depend on the underlying encoder's embedding space. An encoder's training task dictates which acoustic features are preserved or…

音频与语音处理 · 电气工程与系统科学 2026-03-02 Wonwoo Jeong

This paper proposes a method for estimating a surface that contains a given set of points from noisy measurements. More precisely, by assuming that the surface is described by the zero set of a function in the span of a given set of…

系统与控制 · 电气工程与系统科学 2026-04-07 Omar M. Sleem , Sahand Kiani , Constantino M. Lagoa

Head-related transfer functions (HRTFs) describe the directional filtering of the incoming sound caused by the morphology of a listener's head and pinnae. When an accurate model of a listener's morphology exists, HRTFs can be calculated…

数值分析 · 数学 2016-07-28 Harald Ziegelwanger , Wolfgang Kreuzer , Piotr Majdak

Sound field decomposition predicts waveforms in arbitrary directions using signals from a limited number of microphones as inputs. Sound field decomposition is fundamental to downstream tasks, including source localization, source…

声音 · 计算机科学 2022-10-25 Qiuqiang Kong , Shilei Liu , Junjie Shi , Xuzhou Ye , Yin Cao , Qiaoxi Zhu , Yong Xu , Yuxuan Wang

We propose a new robust distributed linearly constrained beamformer which utilizes a set of linear equality constraints to reduce the cross power spectral density matrix to a block-diagonal form. The proposed beamformer has a convenient…

信号处理 · 电气工程与系统科学 2019-05-28 Andreas I. Koutrouvelis , Thomas W. Sherson , Richard Heusdens , Richard C. Hendriks

A method for sound field decomposition based on neural networks is proposed. The method comprises two stages: a sound field separation stage and a single-source localization stage. In the first stage, the sound pressure at microphones…

音频与语音处理 · 电气工程与系统科学 2023-09-14 Ryo Matsuda , Makoto Otani

We focus on automatic feature extraction for raw audio heartbeat sounds, aimed at anomaly detection applications in healthcare. We learn features with the help of an autoencoder composed by a 1D non-causal convolutional encoder and a…

声音 · 计算机科学 2021-02-25 Robert-George Colt , Csongor-Huba Várady , Riccardo Volpi , Luigi Malagò

As spatial audio is enjoying a surge in popularity, data-driven machine learning techniques that have been proven successful in other domains are increasingly used to process head-related transfer function measurements. However, these…

音频与语音处理 · 电气工程与系统科学 2022-12-09 Johan Pauwels , Lorenzo Picinali

We introduce Multi-level feature Fusion-based Periodicity Analysis Model (MF-PAM), a novel deep learning-based pitch estimation model that accurately estimates pitch trajectory in noisy and reverberant acoustic environments. Our model…

音频与语音处理 · 电气工程与系统科学 2025-09-11 Woo-Jin Chung , Doyeon Kim , Soo-Whan Chung , Hong-Goo Kang

This paper proposes to use low-level spatial features extracted from multichannel audio for sound event detection. We extend the convolutional recurrent neural network to handle more than one type of these multichannel features by learning…

声音 · 计算机科学 2017-06-09 Sharath Adavanne , Pasi Pertilä , Tuomas Virtanen