中文
相关论文

相关论文: Explicit Context-Driven Neural Acoustic Modeling f…

200 篇论文

Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings. We propose an audio spatialization method that draws on visual information in videos to convert…

计算机视觉与模式识别 · 计算机科学 2021-11-23 Rishabh Garg , Ruohan Gao , Kristen Grauman

Neural audio synthesis methods can achieve high-fidelity and realistic sound generation by utilizing deep generative models. Such models typically rely on external labels which are often discrete as conditioning information to achieve…

声音 · 计算机科学 2024-06-12 Yunyi Liu , Craig Jin

Recent neural room impulse response (RIR) estimators typically comprise an encoder for reference audio analysis and a generator for RIR synthesis. Especially, it is the performance of the generator that directly influences the overall…

声音 · 计算机科学 2023-11-07 Sungho Lee , Hyeong-Seok Choi , Kyogu Lee

Neural Radiance Field (NeRF) models are implicit neural scene representation methods that offer unprecedented capabilities in novel view synthesis. Semantically-aware NeRFs not only capture the shape and radiance of a scene, but also encode…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Yuzhe Zhu , Lile Cai , Kangkang Lu , Fayao Liu , Xulei Yang

The training of modern speech processing systems often requires a large amount of simulated room impulse response (RIR) data in order to allow the systems to generalize well in real-world, reverberant environments. However, simulating…

音频与语音处理 · 电气工程与系统科学 2022-08-09 Yi Luo , Jianwei Yu

Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: the scarcity of full-band RIR datasets and the inability of…

声音 · 计算机科学 2025-10-30 Ali Vosoughi , Yongyi Zang , Qihui Yang , Nathan Paek , Randal Leistikow , Chenliang Xu

Recently, ray tracing has gained renewed interest with the advent of Reflective Intelligent Surfaces (RIS) technology, a key enabler of 6G wireless communications due to its capability of intelligent manipulation of electromagnetic waves.…

信息论 · 计算机科学 2024-11-07 Huiying Yang , Zihan Jin , Chenhao Wu , Rujing Xiong , Robert Caiming Qiu , Zenan Ling

Implicit Neural Representations (INRs) are widely used for modeling continuous 2D images, enabling high-fidelity reconstruction, super-resolution, and compression. Architectures such as SIREN, WIRE, and FINER demonstrate their ability to…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Weronika Jakubowska , Mikołaj Zieliński , Rafał Tobiasz , Krzysztof Byrski , Maciej Zięba , Dominik Belter , Przemysław Spurek

Implicit Neural Representations have gained prominence as a powerful framework for capturing complex data modalities, encompassing a wide range from 3D shapes to images and audio. Within the realm of 3D shape representation, Neural Signed…

计算机视觉与模式识别 · 计算机科学 2024-08-28 Amine Ouasfi , Adnane Boukhayma

In this work we target a learnable output representation that allows continuous, high resolution outputs of arbitrary shape. Recent works represent 3D surfaces implicitly with a Neural Network, thereby breaking previous barriers in…

计算机视觉与模式识别 · 计算机科学 2020-10-28 Julian Chibane , Aymen Mir , Gerard Pons-Moll

High-dimensional spatio-temporal dynamics can often be encoded in a low-dimensional subspace. Engineering applications for modeling, characterization, design, and control of such large-scale systems often rely on dimensionality reduction to…

机器学习 · 计算机科学 2023-01-05 Shaowu Pan , Steven L. Brunton , J. Nathan Kutz

The approximation and convergence properties of implicit neural representations (INRs) are known to be highly sensitive to parameter initialization strategies. While several data-driven initialization methods demonstrate significant…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Kushal Vyas , Alper Kayabasi , Daniel Kim , Vishwanath Saragadam , Ashok Veeraraghavan , Guha Balakrishnan

Speech audio quality is subject to degradation caused by an acoustic environment and isotropic ambient and point noises. The environment can lead to decreased speech intelligibility and loss of focus and attention by the listener. Basic…

音频与语音处理 · 电气工程与系统科学 2022-04-05 Paula Sánchez López , Paul Callens , Milos Cernak

Recent history has seen a tremendous growth of work exploring implicit representations of geometry and radiance, popularized through Neural Radiance Fields (NeRF). Such works are fundamentally based on a (implicit) volumetric representation…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Jason Y. Zhang , Gengshan Yang , Shubham Tulsiani , Deva Ramanan

Self-supervised learning methods like masked autoencoders (MAE) have shown significant promise in learning robust feature representations, particularly in image reconstruction-based pretraining task. However, their performance is often…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Sua Lee , Joonhun Lee , Myungjoo Kang

Neural Radiance Field (NeRF) has revolutionized novel-view rendering tasks and achieved impressive results. However, the inefficient sampling and per-scene optimization hinder its wide applications. Though some generalizable NeRFs have been…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Yue Shi , Dingyi Rong , Chang Chen , Chaofan Ma , Bingbing Ni , Wenjun Zhang

Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have…

音频与语音处理 · 电气工程与系统科学 2024-02-29 Yoshiki Masuyama , Gordon Wichern , François G. Germain , Zexu Pan , Sameer Khurana , Chiori Hori , Jonathan Le Roux

This paper tackles the problem of novel view audio-visual synthesis along an arbitrary trajectory in an indoor scene, given the audio-video recordings from other known trajectories of the scene. Existing methods often overlook the effect of…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Huiyu Gao , Jiahao Ma , David Ahmedt-Aristizabal , Chuong Nguyen , Miaomiao Liu

While deep learning reshaped the classical motion capture pipeline with feed-forward networks, generative models are required to recover fine alignment via iterative refinement. Unfortunately, the existing models are usually hand-crafted or…

计算机视觉与模式识别 · 计算机科学 2021-11-01 Shih-Yang Su , Frank Yu , Michael Zollhoefer , Helge Rhodin

Neural implicit surface representations have emerged as a promising paradigm to capture 3D shapes in a continuous and resolution-independent manner. However, adapting them to articulated shapes is non-trivial. Existing approaches learn a…

计算机视觉与模式识别 · 计算机科学 2021-11-30 Xu Chen , Yufeng Zheng , Michael J. Black , Otmar Hilliges , Andreas Geiger