中文
相关论文

相关论文: Reciprocal Latent Fields for Precomputed Sound Pro…

200 篇论文

Leveraging representation encoders for generative modeling offers a path for efficient, high-fidelity synthesis. However, standard diffusion transformers fail to converge on these representations directly. While recent work attributes this…

机器学习 · 计算机科学 2026-02-11 Amandeep Kumar , Vishal M. Patel

Due to strict rate and reliability demands, wireless image transmission remains difficult for both classical layered designs and joint source-channel coding (JSCC), especially under low latency. Diffusion-based generative decoders can…

机器学习 · 计算机科学 2026-01-13 Jingwen Fu , Ming Xiao , Mikael Skoglund , Dong In Kim

Reinforcement learning (RL) offers a capable and intuitive structure for the fundamental sequential decision-making problem. Despite impressive breakthroughs, it can still be difficult to employ RL in practice in many simple applications.…

人工智能 · 计算机科学 2024-01-18 Aida Afshar , Wenchao Li

Deep Neural Networks are known to be very demanding in terms of computing and memory requirements. Due to the ever increasing use of embedded systems and mobile devices with a limited resource budget, designing low-complexity models without…

机器学习 · 计算机科学 2020-11-06 Khaled Koutini , Florian Henkel , Hamid Eghbal-zadeh , Gerhard Widmer

Generative models have achieved remarkable progress with the emergence of flow matching (FM). It has demonstrated strong generative capabilities and attracted significant attention as a simulation-free flow-based framework capable of…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Huynh Trinh Ngoc , Hoang Anh Nguyen Kim , Toan Nguyen Hai , Long Tran Quoc

This paper introduces the Procedural (audio) Variational autoEncoder (ProVE) framework as a general approach to learning Procedural Audio PA models of environmental sounds with an improvement to the realism of the synthesis while…

声音 · 计算机科学 2023-03-07 Danzel Serrano , Mark Cartwright

For 6-DOF (degrees of freedom) interactive virtual acoustic environments (VAEs), the spatial rendering of diffuse late reverberation in addition to early (specular) reflections is important. In the interest of computational efficiency, the…

音频与语音处理 · 电气工程与系统科学 2021-11-30 Christoph Kirsch , Josef Poppitz , Torben Wendt , Steven van de Par , Stephan D. Ewert

Recently, ray tracing has gained renewed interest with the advent of Reflective Intelligent Surfaces (RIS) technology, a key enabler of 6G wireless communications due to its capability of intelligent manipulation of electromagnetic waves.…

信息论 · 计算机科学 2024-11-07 Huiying Yang , Zihan Jin , Chenhao Wu , Rujing Xiong , Robert Caiming Qiu , Zenan Ling

Addressing the critical need for robust safety in Large Language Models (LLMs), particularly against adversarial attacks and in-distribution errors, we introduce Reinforcement Learning with Backtracking Feedback (RLBF). This framework…

机器学习 · 计算机科学 2026-04-28 Bilgehan Sel , Vaishakh Keshava , Phillip Wallis , Lukas Rutishauser , Ming Jin , Dingcheng Li

The recent work Local Implicit Image Function (LIIF) and subsequent Implicit Neural Representation (INR) based works have achieved remarkable success in Arbitrary-Scale Super-Resolution (ASSR) by using MLP to decode Low-Resolution (LR)…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Zongyao He , Zhi Jin

Understanding how explicit theoretical features are encoded in opaque neural systems is a central challenge now common to neuroscience and AI. We introduce Metric Learning Encoding Models (MLEMs) to address this challenge most directly as a…

计算与语言 · 计算机科学 2025-11-17 Louis Jalouzot , Christophe Pallier , Emmanuel Chemla , Yair Lakretz

Neural Radiance Fields (NeRF) is a popular view synthesis technique that represents a scene as a continuous volumetric function, parameterized by multilayer perceptrons that provide the volume density and view-dependent emitted radiance at…

计算机视觉与模式识别 · 计算机科学 2021-12-08 Dor Verbin , Peter Hedman , Ben Mildenhall , Todd Zickler , Jonathan T. Barron , Pratul P. Srinivasan

Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA)…

声音 · 计算机科学 2024-03-22 Samuel Pegg , Kai Li , Xiaolin Hu

Personalized Head-Related Transfer Functions (HRTFs) are starting to be introduced in many commercial immersive audio applications and are crucial for realistic spatial audio rendering. However, one of the main hesitations regarding their…

声音 · 计算机科学 2025-10-03 Xuyi Hu , Jian Li , Shaojie Zhang , Stefan Goetz , Lorenzo Picinali , Ozgur B. Akan , Aidan O. T. Hogg

Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial upsampling has shown remarkable progress with neural fields,…

音频与语音处理 · 电气工程与系统科学 2025-01-23 Yoshiki Masuyama , Gordon Wichern , François G. Germain , Christopher Ick , Jonathan Le Roux

Most of the prevalent approaches in speech prosody modeling rely on learning global style representations in a continuous latent space which encode and transfer the attributes of reference speech. However, recent work on neural codecs which…

Latent force models (LFM) are principled approaches to incorporating solutions to differential equations within non-parametric inference methods. Unfortunately, the development and application of LFMs can be inhibited by their computational…

机器学习 · 统计学 2014-05-30 Steven Reece , Stephen Roberts , Siddhartha Ghosh , Alex Rogers , Nicholas Jennings

This paper introduces a novel framework for image quality transfer based on conditional flow matching (CFM). Unlike conventional generative models that rely on iterative sampling or adversarial objectives, CFM learns a continuous flow…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Huu Tien Nguyen , Ahmed Karam Eldaly

GeRaF is the first method to use neural implicit learning for near-range 3D geometry reconstruction from radio frequency (RF) signals. Unlike RGB or LiDAR-based methods, RF sensing can see through occlusion but suffers from low resolution…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Jiachen Lu , Hailan Shanbhag , Haitham Al Hassanieh

Frame stacking is broadly applied in end-to-end neural network training like connectionist temporal classification (CTC), and it leads to more accurate models and faster decoding. However, it is not well-suited to conventional neural…

计算与语言 · 计算机科学 2017-05-18 Xu Tian , Jun Zhang , Zejun Ma , Yi He , Juan Wei