中文
相关论文

相关论文: Room Impulse Response Generation Conditioned on Ac…

200 篇论文

Discrete diffusion models generate sequences by iteratively denoising samples corrupted by categorical noise, offering an appealing alternative to autoregressive decoding for structured and symbolic generation. However, standard training…

机器学习 · 计算机科学 2026-02-04 Huu Binh Ta , Michael Cardei , Alvaro Velasquez , Ferdinando Fioretto

A Recurrent Neural Network (RNN) for audio synthesis is trained by augmenting the audio input with information about signal characteristics such as pitch, amplitude, and instrument. The result after training is an audio synthesizer that is…

声音 · 计算机科学 2018-05-31 Lonce Wyse

Speech super-resolution (SR) is the task that restores high-resolution speech from low-resolution input. Existing models employ simulated data and constrained experimental settings, which limit generalization to real-world SR. Predictive…

音频与语音处理 · 电气工程与系统科学 2024-01-26 Heming Wang , Eric W. Healy , DeLiang Wang

Realistic music generation is a challenging task. When building generative models of music that are learnt from data, typically high-level representations such as scores or MIDI are used that abstract away the idiosyncrasies of a particular…

声音 · 计算机科学 2018-06-28 Sander Dieleman , Aäron van den Oord , Karen Simonyan

Acoustic Environment Matching (AEM) is the task of transferring clean audio into a target acoustic environment, enabling engaging applications such as audio dubbing and auditory immersive virtual reality (VR). Recovering similar room…

声音 · 计算机科学 2026-04-01 Chenpei Huang , Lingfeng Yao , Kyu In Lee , Lan Emily Zhang , Xun Chen , Miao Pan

In this work, we introduce a novel framework which combines physics and machine learning methods to analyse acoustic signals. Three methods are developed for this task: a Bayesian inference approach for inferring the spectral acoustics…

声音 · 计算机科学 2023-05-30 Yongchao Huang , Yuhang He , Hong Ge

Implicit neural representations (INRs) have emerged as a promising approach for video storage and processing, showing remarkable versatility across various video tasks. However, existing methods often fail to fully leverage their…

图像与视频处理 · 电气工程与系统科学 2024-03-19 Xinjie Zhang , Ren Yang , Dailan He , Xingtong Ge , Tongda Xu , Yan Wang , Hongwei Qin , Jun Zhang

Predicting Room Impulse Responses (RIRs) remains a challenge due to the high dimensionality of audio signals and the need for perceptual accuracy. This paper introduces a neural network framework that predicts multi-band Energy Decay Curves…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Imran Muhammad , Gerald Schuller

We introduce a novel algorithm for online estimation of acoustic impulse responses (AIRs) which allows for fast convergence by exploiting prior knowledge about the fundamental structure of AIRs. The proposed method assumes that the…

音频与语音处理 · 电气工程与系统科学 2021-05-10 Thomas Haubner , Andreas Brendel , Walter Kellermann

Predicting spatially varying Room Impulse Response (RIR) from sparse observations is a critical but highly challenging inverse problem for immersive spatial audio rendering. In this work, we present EIGENET, a geometry-informed multi-modal…

声音 · 计算机科学 2026-05-28 Chong Jing , Zitong Lan , Junan Zhang , Zhizheng Wu

Speech enhancement in hearing aids remains a difficult task in nonstationary acoustic environments, mainly because current signal processing algorithms rely on fixed, manually tuned parameters that cannot adapt in situ to different users or…

The image-source method is widely applied to compute room impulse responses (RIRs) of shoebox rooms with arbitrary absorption. However, with increasing RIR lengths, the number of image sources grows rapidly, leading to slow computation. In…

音频与语音处理 · 电气工程与系统科学 2023-10-12 Sebastian J. Schlecht , Karolina Prawda , Rudolf Rabenstein , Maximilian Schäfer

Generating sound effects that humans want is an important topic. However, there are few studies in this area for sound generation. In this study, we investigate generating sound conditioned on a text prompt and propose a novel text-to-sound…

声音 · 计算机科学 2023-05-01 Dongchao Yang , Jianwei Yu , Helin Wang , Wen Wang , Chao Weng , Yuexian Zou , Dong Yu

Diffusion models have recently been shown to be relevant for high-quality speech generation. Most work has been focused on generating spectrograms, and as such, they further require a subsequent model to convert the spectrogram to a…

声音 · 计算机科学 2024-03-12 Roi Benita , Michael Elad , Joseph Keshet

Existing Reward Models (RMs), typically trained on general preference data, struggle in Retrieval Augmented Generation (RAG) settings, which require judging responses for faithfulness to retrieved context, relevance to the user query,…

Generative recommendation (GR) has emerged as a promising paradigm that predicts target items by autoregressively generating their semantic identifiers (SID). Most GR methods follow a quantization-representation-generation pipeline, first…

信息检索 · 计算机科学 2026-05-13 Ziwei Liu , Yejing Wang , Shengyu Zhou , Xinhang Li , Xiangyu Zhao

Deep generative models have emerged as a promising approach in the medical image domain to address data scarcity. However, their use for sequential data like respiratory sounds is less explored. In this work, we propose a straightforward…

声音 · 计算机科学 2023-11-14 June-Woo Kim , Chihyeon Yoon , Miika Toikkanen , Sangmin Bae , Ho-Young Jung

Reconfigurable intelligent surface (RIS) is a promising technique to enhance the performance of physical-layer key generation (PKG) due to its ability to smartly customize the radio environments. Existing RIS-assisted PKG methods are mainly…

信息论 · 计算机科学 2022-07-26 Lei Hu , Guyue Li , Xuewen Qian , Derrick Wing Kwan Ng , Aiqun Hu

Our everyday auditory experience is shaped by the acoustics of the indoor environments in which we live. Room acoustics modeling is aimed at establishing mathematical representations of acoustic wave propagation in such environments. These…

音频与语音处理 · 电气工程与系统科学 2025-04-24 Toon van Waterschoot

Practically, training diffusion models typically requires explicit time conditioning to guide the network through the denoising sampling process. Especially in deterministic methods like DDIM, the absence of time conditioning leads to…

机器学习 · 计算机科学 2026-04-29 Liuzhuozheng Li , Zhiyuan Zhan , Shuhong Liu , Dengyang Jiang , Zanyi Wang , Guang Dai , Jingdong Wang , Mengmeng Wang