中文
相关论文

相关论文: Learning Magnitude Distribution of Sound Fields vi…

200 篇论文

Personalized Head-Related Transfer Functions (HRTFs) are starting to be introduced in many commercial immersive audio applications and are crucial for realistic spatial audio rendering. However, one of the main hesitations regarding their…

声音 · 计算机科学 2025-10-03 Xuyi Hu , Jian Li , Shaojie Zhang , Stefan Goetz , Lorenzo Picinali , Ozgur B. Akan , Aidan O. T. Hogg

Neural audio autoencoders create compact latent representations that preserve perceptually important information, serving as the foundation for both modern audio compression systems and generation approaches like next-token prediction and…

声音 · 计算机科学 2025-09-10 Dimitrios Bralios , Paris Smaragdis , Jonah Casebeer

This paper addresses the problem of binaural localization of a single speech source in noisy and reverberant environments. For a given binaural microphone setup, the binaural response corresponding to the direct-path propagation of a single…

声音 · 计算机科学 2016-09-08 Xiaofei Li , Laurent Girin , Radu Horaud , Sharon Gannot

A novel approach for speech segmentation is proposed, based on Multilevel Hybrid (mean/min) Filters (MHF) with the following features: An accurate transition location. Good performance in noisy environments (gaussian and impulsive noise).…

音频与语音处理 · 电气工程与系统科学 2022-03-04 Marcos Faundez-Zanuy , Francesc Vallverdu-Bayes

Accurate and reliable identification of the relative transfer functions (RTFs) between microphones with respect to a desired source is an essential component in the design of microphone array beamformers, specifically when applying the…

音频与语音处理 · 电气工程与系统科学 2024-12-19 Daniel Levi , Amit Sofer , Sharon Gannot

This paper proposes a blind estimation method based on the modulation transfer function and Schroeder model for estimating reverberation time in seven-octave bands. Therefore, the speech transmission index and five room-acoustic parameters…

声音 · 计算机科学 2021-03-16 Suradej Duangpummet , Jessada Karnjana , Waree Kongprawechnon , Masashi Unoki

While there has been much recent progress using deep learning techniques to separate speech and music audio signals, these systems typically require large collections of isolated sources during the training process. When extending audio…

声音 · 计算机科学 2020-09-01 Fatemeh Pishdadian , Gordon Wichern , Jonathan Le Roux

\textbf{Purpose:} Amplitude analysis is a pivotal tool in hadron spectroscopy, fundamentally involving a series of likelihood fits to multi-dimensional experimental distributions. While robust goodness-of-fit tests exist for low-dimensional…

数据分析、统计与概率 · 物理学 2025-12-02 Huoyi Hou , Beijiang Liu

In neural audio signal processing, pitch conditioning has been used to enhance the performance of synthesizers. However, jointly training pitch estimators and synthesizers is a challenge when using standard audio-to-audio reconstruction…

声音 · 计算机科学 2024-01-17 Bernardo Torres , Geoffroy Peeters , Gaël Richard

Traditional sound diffusers are quasi-random phase gratings attached to reflecting surfaces whose purpose is to augment the spatiotemporal incoherence of the acoustic field scattered from reflective surfaces. This configuration allows one…

应用物理 · 物理学 2022-11-23 Janghoon Kang , Michael R. Haberman

We describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a speaker identification task. The learned STRFs were then…

音频与语音处理 · 电气工程与系统科学 2021-02-16 Tyler Vuong , Yangyang Xia , Richard M. Stern

Acoustic scene classification (ASC) aims to identify the type of scene (environment) in which a given audio signal is recorded. The log-mel feature and convolutional neural network (CNN) have recently become the most popular time-frequency…

声音 · 计算机科学 2021-08-12 Yuzhong Wu , Tan Lee

We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed for the generative…

声音 · 计算机科学 2023-10-27 Ali Golmakani , Mostafa Sadeghi , Xavier Alameda-Pineda , Romain Serizel

Sound field reconstruction refers to the problem of estimating the acoustic pressure field over an arbitrary region of space, using only a limited set of measurements. Physics-informed neural networks have been adopted to solve the problem…

音频与语音处理 · 电气工程与系统科学 2025-06-05 Stefano Damiano , Toon van Waterschoot

In industrial applications, the early detection of malfunctioning factory machinery is crucial. In this paper, we consider acoustic malfunction detection via transfer learning. Contrary to the majority of current approaches which are based…

音频与语音处理 · 电气工程与系统科学 2021-02-19 Robert Müller , Fabian Ritz , Steffen Illium , Claudia Linnhoff-Popien

Accurate knowledge of acoustic surface admittance or impedance is essential for reliable wave-based simulations, yet its in situ estimation remains challenging due to noise, model inaccuracies, and restrictive assumptions of conventional…

机器学习 · 计算机科学 2026-04-10 Jonas M. Schmid , Johannes D. Schmid , Martin Eser , Steffen Marburg

Maximum Voiced Frequency (MVF) is used in various speech models as the spectral boundary separating periodic and aperiodic components during the production of voiced sounds. Recent studies have shown that its proper estimation and modeling…

音频与语音处理 · 电气工程与系统科学 2020-06-02 Thomas Drugman , Yannis Stylianou

In many multi-microphone algorithms for noise reduction, an estimate of the relative transfer function (RTF) vector of the target speaker is required. The state-of-the-art covariance whitening (CW) method estimates the RTF vector as the…

音频与语音处理 · 电气工程与系统科学 2023-10-30 Wiebke Middelberg , Henri Gode , Simon Doclo

An interpolation method for region-to-region acoustic transfer functions (ATFs) based on kernel ridge regression with an adaptive kernel is proposed. Most current ATF interpolation methods do not incorporate the acoustic properties for…

音频与语音处理 · 电气工程与系统科学 2023-03-08 Juliano G. C. Ribeiro , Shoichi Koyama , Hiroshi Saruwatari

Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial upsampling has shown remarkable progress with neural fields,…

音频与语音处理 · 电气工程与系统科学 2025-01-23 Yoshiki Masuyama , Gordon Wichern , François G. Germain , Christopher Ick , Jonathan Le Roux