中文
相关论文

相关论文: Beyond Omnidirectional: Neural Ambisonics Encoding…

200 篇论文

In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech enhancement system remains challenging, especially in…

音频与语音处理 · 电气工程与系统科学 2024-12-31 Wen Wen , Qiang Zhou , Yu Xi , Haoyu Li , Ziqi Gong , Kai Yu

Spatial audio methods are gaining a growing interest due to the spread of immersive audio experiences and applications, such as virtual and augmented reality. For these purposes, 3D audio signals are often acquired through arrays of…

音频与语音处理 · 电气工程与系统科学 2022-12-16 Eleonora Grassucci , Gioia Mancini , Christian Brignone , Aurelio Uncini , Danilo Comminiello

Deep learning dominates speech processing but relies on massive datasets, global backpropagation-guided weight updates, and produces entangled representations. Assembly Calculus (AC), which models sparse neuronal assemblies via Hebbian…

音频与语音处理 · 电气工程与系统科学 2026-03-19 Trevor Adelson , Vidhyasaharan Sethu , Ting Dang

Latest advances in deep spatial filtering for Ambisonics demonstrate strong performance in stationary multi-speaker scenarios by rotating the sound field toward a target speaker prior to multi-channel enhancement. For applicability in…

音频与语音处理 · 电气工程与系统科学 2026-01-22 Jakob Kienegger , Timo Gerkmann

Approaches to signal representation and coding theory have traditionally focused on how to best represent signals using parsimonious representations that incur the lowest possible distortion. Classical examples include linear and non-linear…

信息论 · 计算机科学 2015-12-25 Petros T Boufounos , Shantanu Rane , Hassan Mansour

Optomechanical arrays are a promising future platform for studies of transport, many-body dynamics, quantum control and topological effects in systems of coupled photon and phonon modes. We introduce disordered optomechanical arrays,…

介观与纳米尺度物理 · 物理学 2017-09-15 Thales Figueiredo Roque , Vittorio Peano , Oleg M. Yevtushenko , Florian Marquardt

Sound event localization frameworks based on deep neural networks have shown increased robustness with respect to reverberation and noise in comparison to classical parametric approaches. In particular, recurrent architectures that…

Audio zooming, a signal processing technique, enables selective focusing and enhancement of sound signals from a specified region, attenuating others. While traditional beamforming and neural beamforming techniques, centered on creating a…

音频与语音处理 · 电气工程与系统科学 2023-11-23 Meng Yu , Dong Yu

Active noise control (ANC) over a sizeable space requires a large number of reference and error microphones to satisfy the spatial Nyquist sampling criterion, which limits the feasibility of practical realization of such systems. This paper…

声音 · 计算机科学 2018-03-02 Yu Maeno , Yuki Mitsufuji , Thushara D. Abhayapala

Invariance to microphone array configuration is a rare attribute in neural beamformers. Filter-and-sum (FS) methods in this class define the target signal with respect to a reference channel. However, this not only complicates formulation…

音频与语音处理 · 电气工程与系统科学 2023-02-28 Anton Kovalyov , Kashyap Patel , Issa Panahi

Conventional direction of arrival (DOA) estimators are based on array processing using either time differences or beamforming. The proposed approach is based on the received power at each microphone, which enables simple hardware, low…

信号处理 · 电气工程与系统科学 2025-09-16 Gustav Zetterqvist , Fredrik Gustafsson , Gustaf Hendeby

Rigid spherical microphone arrays (RSMAs) have been widely used in ambisonics sound field recording. While it is desired to combine the information captured by a grid of densely arranged RSMAs for expanding the area of accurate…

声音 · 计算机科学 2021-08-17 Shoken Kaneko , Ramani Duraiswami

In this work, we propose a new spatio-directional neural encoding that is compact and efficient, and supports all-frequency signals in both space and direction. Current learnable encodings focus on Cartesian orthonormal spaces, which have…

图形学 · 计算机科学 2026-03-06 Philippe Weier , Lukas Bode , Philipp Slusallek , Adrián Jarabo , Sébastien Speierer

Neural networks (NNs) have been widely applied in speech processing tasks, and, in particular, those employing microphone arrays. Nevertheless, most existing NN architectures can only deal with fixed and position-specific microphone arrays.…

音频与语音处理 · 电气工程与系统科学 2021-06-14 Yochai Yemini , Ethan Fetaya , Haggai Maron , Sharon Gannot

Augmented listening devices, such as hearing aids and augmented reality headsets, enhance human perception by changing the sounds that we hear. Microphone arrays can improve the performance of listening systems in noisy environments, but…

音频与语音处理 · 电气工程与系统科学 2020-04-28 Ryan M. Corey , Andrew C. Singer

Multichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions to the real-world. Distributed sensor arrays that consider several devices with a few microphones is a viable…

声音 · 计算机科学 2020-03-17 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

Audio coding is an essential module in the real-time communication system. Neural audio codecs can compress audio samples with a low bitrate due to the strong modeling and generative capabilities of deep neural networks. To address the poor…

声音 · 计算机科学 2023-10-18 Wenzhe Liu , Wei Xiao , Meng Wang , Shan Yang , Yupeng Shi , Yuyong Kang , Dan Su , Shidong Shang , Dong Yu

Arbitrary linear transformations are of crucial importance in a plethora of photonic applications spanning classical signal processing, communication systems, quantum information processing and machine learning. Here, we present a new…

This study presents a deep-learning framework for controlling multichannel acoustic feedback in audio devices. Traditional digital signal processing methods struggle with convergence when dealing with highly correlated noise such as…

声音 · 计算机科学 2025-05-30 Yuan-Kuei Wu , Juan Azcarreta , Kashyap Patel , Buye Xu , Jung-Suk Lee , Sanha Lee , Ashutosh Pandey

We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress…