中文
相关论文

相关论文: NIIRF: Neural IIR Filter Field for HRTF Upsampling…

200 篇论文

This paper tackles two major problem settings for interpretability of audio processing networks, post-hoc and by-design interpretation. For post-hoc interpretation, we aim to interpret decisions of a network in terms of high-level audio…

Hearing aid is an electroacoustic device used to selectively amplify the audio sounds with an aim to make speech more intelligible for a hearing impaired person. Filter bank is one of the important parts of digital hearing aid where the sub…

声音 · 计算机科学 2020-12-22 Arun Sebastian , Manu Francis , Arun Mathew

In this paper, we propose a novel soft and monotonic alignment mechanism used for sequence transduction. It is inspired by the integrate-and-fire model in spiking neural networks and employed in the encoder-decoder framework consists of…

计算与语言 · 计算机科学 2020-02-13 Linhao Dong , Bo Xu

Neural Fields (NF) have gained prominence as a versatile framework for complex data representation. This work unveils a new problem setting termed \emph{Meta-Continual Learning of Neural Fields} (MCL-NF) and introduces a novel strategy that…

人工智能 · 计算机科学 2026-02-24 Seungyoon Woo , Junhyeog Yun , Gunhee Kim

The Orientation Distribution Function (ODF) characterizes key brain microstructural properties and plays an important role in understanding brain structural connectivity. Recent works introduced Implicit Neural Representation (INR) based…

图像与视频处理 · 电气工程与系统科学 2024-09-17 Mohammed Munzer Dwedari , William Consagra , Philip Müller , Özgün Turgut , Daniel Rueckert , Yogesh Rathi

Many spatial filtering algorithms used for voice capture in, e.g., teleconferencing applications, can benefit from or even rely on knowledge of Relative Transfer Functions (RTFs). Accordingly, many RTF estimators have been proposed which,…

音频与语音处理 · 电气工程与系统科学 2021-10-06 Andreas Brendel , Johannes Zeitler , Walter Kellermann

Precise elevation perception in binaural audio remains a challenge, despite extensive research on head-related transfer functions (HRTFs) and spectral cues. While prior studies have advanced our understanding of sound localization cues, the…

信号处理 · 电气工程与系统科学 2025-03-17 Juan Antonio De Rus , Mario Montagud , Jesus Lopez-Ballester , Francesc J. Ferri , Maximo Cobos

Neural waveform models such as WaveNet have demonstrated better performance than conventional vocoders for statistical parametric speech synthesis. As an autoregressive (AR) model, WaveNet is limited by a slow sequential waveform generation…

音频与语音处理 · 电气工程与系统科学 2019-11-19 Xin Wang , Shinji Takaki , Junichi Yamagishi

Neural Radiance Fields (NeRFs) have emerged as powerful tools for capturing detailed 3D scenes through continuous volumetric representations. Recent NeRFs utilize feature grids to improve rendering quality and speed; however, these…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Tuan Pham , Stephan Mandt

Hinged on the representation power of neural networks, neural radiance fields (NeRF) have recently emerged as one of the promising and widely applicable methods for 3D object and scene representation. However, NeRF faces challenges in…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Jaeyoung Chung , Kanggeon Lee , Sungyong Baik , Kyoung Mu Lee

Neural radiance fields (NeRFs) have become a ubiquitous tool for modeling scene appearance and geometry from multiview imagery. Recent work has also begun to explore how to use additional supervision from lidar or depth sensor measurements…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Anagh Malik , Parsa Mirdehghan , Sotiris Nousias , Kiriakos N. Kutulakos , David B. Lindell

In this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a…

音频与语音处理 · 电气工程与系统科学 2026-02-18 Ilai Zaidel , Sharon Gannot

Automatic speech recognition (ASR) on multi-talker recordings is challenging. Current methods using 3D spatial data from multi-channel audio and visual cues focus mainly on direct waves from the target speaker, overlooking reflection wave…

音频与语音处理 · 电气工程与系统科学 2024-06-13 Yiwen Shao , Shi-Xiong Zhang , Dong Yu

The multichannel Wiener filter (MWF) and its variations have been extensively applied to binaural hearing aids. However, its major drawback is the distortion of the binaural cues of the residual noise, changing the original acoustic…

音频与语音处理 · 电气工程与系统科学 2019-11-12 Johnny Werner , Marcio H. Costa

An inverse source reconstruction (ISR) based 3-D near-field (NF) passive radar microwave imaging method utilizing modulated signals is presented. The modulated signals from a non-cooperative transmitter are scattered by the targets of…

图像与视频处理 · 电气工程与系统科学 2026-05-06 Quanfeng Wang , Alexander H. Paulus , Thomas F. Eibert

Largely due to their implicit nature, neural fields lack a direct mechanism for filtering, as Fourier analysis from discrete signal processing is not directly applicable to these representations. Effective filtering of neural fields is…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Ahan Shabanov , Shrisudhan Govindarajan , Cody Reading , Lily Goli , Daniel Rebain , Kwang Moo Yi , Andrea Tagliasacchi

Speaker localization for binaural microphone arrays has been widely studied for applications such as speech communication, video conferencing, and robot audition. Many methods developed for this task, including the direct path dominance…

音频与语音处理 · 电气工程与系统科学 2023-11-01 Yanir Maymon , Israel Nelken , Boaz Rafaely

Utilizing spherical harmonic (SH) domain has been established as the default method of obtaining continuity over space in head-related transfer functions (HRTFs). This paper concerns different variants of extending this solution by…

音频与语音处理 · 电气工程与系统科学 2023-07-19 Adam Szwajcowski

Neural Radiance Fields (NeRF) are able to reconstruct scenes with unprecedented fidelity, and various recent works have extended NeRF to handle dynamic scenes. A common approach to reconstruct such non-rigid scenes is through the use of a…

计算机视觉与模式识别 · 计算机科学 2021-09-13 Keunhong Park , Utkarsh Sinha , Peter Hedman , Jonathan T. Barron , Sofien Bouaziz , Dan B Goldman , Ricardo Martin-Brualla , Steven M. Seitz

Conventional static measurement of head-related impulse responses (HRIRs) is time-consuming due to the need for repositioning a speaker array for each azimuth angle. Dynamic approaches using analytical models with a continuously rotating…

音频与语音处理 · 电气工程与系统科学 2025-04-22 Byeong-Yun Ko , Deokki Min , Hyeonuk Nam , Yong-Hwa Park