中文
相关论文

相关论文: Learning Magnitude Distribution of Sound Fields vi…

200 篇论文

Reconstructing the room transfer functions needed to calculate the complex sound field in a room has several important real-world applications. However, an unpractical number of microphones is often required. Recently, in addition to…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Francesca Ronchini , Luca Comanducci , Mirco Pezzoli , Fabio Antonacci , Augusto Sarti

The signal to noise ratio (SNR) is one of the important measures for reducing the noise.A technique that uses a linear prediction error filter (LPEF) and an adaptive digital filter (ADF) to achieve noise reduction in a speech and image…

网络与互联网体系结构 · 计算机科学 2011-10-12 R. Seshadri , N. Penchalaiah

Split-Federated (SplitFed) learning is an extension of federated learning that places minimal requirements on the clients computing infrastructure, since only a small portion of the overall model is deployed on the clients hardware. In…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Zahra Hafezi Kafshgari , Ivan V. Bajic , Parvaneh Saeedi

We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. F0 estimation is important in various speech processing and music…

音频与语音处理 · 电气工程与系统科学 2021-02-15 Satwinder Singh , Ruili Wang , Yuanhang Qiu

Matched filters are widely used to localise signal patterns due to their high efficiency and interpretability. However, their effectiveness deteriorates for low signal-to-noise ratio (SNR) signals, such as those recorded on edge devices,…

信号处理 · 电气工程与系统科学 2025-09-01 Haozhe Tian , Qiyu Rao , Nina Moutonnet , Pietro Ferraro , Danilo Mandic

Automatic Music Transcription (AMT) -- the task of converting music audio into note representations -- has seen rapid progress, driven largely by deep learning systems. Due to the limited availability of richly annotated music datasets,…

声音 · 计算机科学 2026-01-27 Lukáš Samuel Marták , Patricia Hu , Gerhard Widmer

Accurate acoustic simulations of enclosed spaces require precise boundary conditions, typically expressed through surface impedances for wave-based methods. Conventional measurement techniques often rely on simplifying assumptions about the…

声音 · 计算机科学 2026-04-09 Jonas M. Schmid , Johannes D. Schmid , Martin Eser , Steffen Marburg

The modulation transfer function (MTF) represents the frequency domain response of imaging modalities. Here, we report a method for estimating the MTF from sample images. Test images were generated from a number of images, including those…

图像与视频处理 · 电气工程与系统科学 2017-12-05 Rino Saiga , Akihisa Takeuchi , Kentaro Uesugi , Yasuko Terada , Yoshio Suzuki , Ryuta Mizutani

Recently, hyperspherical embeddings have established themselves as a dominant technique for face and voice recognition. Specifically, Euclidean space vector embeddings are learned to encode person-specific information in their direction…

音频与语音处理 · 电气工程与系统科学 2022-10-25 Nikita Kuzmin , Igor Fedorov , Alexey Sholokhov

In this paper, a deep-learning-based method for sound field reconstruction is proposed. It is shown the possibility to reconstruct the magnitude of the sound pressure in the frequency band 30-300 Hz for an entire room by using a very low…

声音 · 计算机科学 2020-08-07 Francesc Lluís , Pablo Martínez-Nuevo , Martin Bo Møller , Sven Ewan Shepstone

Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two microphone channels. Though DP-RTF fully encodes the sound spatial cues and serves as a reliable localization…

声音 · 计算机科学 2022-02-17 Bing Yang , Hong Liu , Xiaofei Li

Active noise control (ANC) over a sizeable space requires a large number of reference and error microphones to satisfy the spatial Nyquist sampling criterion, which limits the feasibility of practical realization of such systems. This paper…

声音 · 计算机科学 2018-03-02 Yu Maeno , Yuki Mitsufuji , Thushara D. Abhayapala

Seismic acoustic impedance plays a crucial role in lithological identification and subsurface structure interpretation. However, due to the inherently ill-posed nature of the inversion problem, directly estimating impedance from post-stack…

机器学习 · 计算机科学 2025-06-17 Jie Chen , Hongling Chen , Jinghuai Gao , Chuangji Meng , Tao Yang , XinXin Liang

This paper studies the fundamental limit of semantic communications over the discrete memoryless channel. We consider the scenario to send a semantic source consisting of an observation state and its corresponding semantic state, both of…

信息论 · 计算机科学 2024-01-03 Dongxu Li , Jianhao Huang , Chuan Huang , Xiaoqi Qin , Han Zhang , Ping Zhang

We learn audio representations by solving a novel self-supervised learning task, which consists of predicting the phase of the short-time Fourier transform from its magnitude. A convolutional encoder is used to map the magnitude spectrum of…

音频与语音处理 · 电气工程与系统科学 2019-10-29 Félix de Chaumont Quitry , Marco Tagliasacchi , Dominik Roblek

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

计算与语言 · 计算机科学 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich

Music emotion recognition (MER), a sub-task of music information retrieval (MIR), has developed rapidly in recent years. However, the learning of affect-salient features remains a challenge. In this paper, we propose an end-to-end…

声音 · 计算机科学 2022-07-01 Zi Huang , Shulei Ji , Zhilan Hu , Chuangjian Cai , Jing Luo , Xinyu Yang

Squeezed vacuum states are now employed in gravitational-wave interferometric detectors, enhancing their sensitivity and thus enabling richer astrophysical observations. In future observing runs, the detectors will incorporate a filter…

天体物理仪器与方法 · 物理学 2022-07-13 Dhruva Ganapathy , Victoria Xu , Wenxuan Jia , Chris Whittle , Maggie Tse , Lisa Barsotti , Matthew Evans , Lee McCuller

Deep learning-based Personal Sound Zones (PSZs) rely on simulated acoustic transfer functions (ATFs) for training, yet idealized point-source models exhibit large sim-to-real gaps. While physically informed components improve…

音频与语音处理 · 电气工程与系统科学 2026-03-04 Hao Jiang , Edgar Choueiri

Conventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling has shown additional improvements by training a spatial…

声音 · 计算机科学 2020-02-10 Taejin Park , Kenichi Kumatani , Minhua Wu , Shiva Sundaram