中文
相关论文

相关论文: On the Usefulness of Diffusion-Based Room Impulse …

200 篇论文

We present a simple and efficient method for refining maps or correspondences by iterative upsampling in the spectral domain that can be implemented in a few lines of code. Our main observation is that high quality maps can be obtained even…

图形学 · 计算机科学 2019-09-13 Simone Melzi , Jing Ren , Emanuele Rodolà , Abhishek Sharma , Peter Wonka , Maks Ovsjanikov

Voice conversion is a common speech synthesis task which can be solved in different ways depending on a particular real-world scenario. The most challenging one often referred to as one-shot many-to-many voice conversion consists in copying…

声音 · 计算机科学 2022-08-05 Vadim Popov , Ivan Vovk , Vladimir Gogoryan , Tasnima Sadekova , Mikhail Kudinov , Jiansheng Wei

While most scene flow methods use either variational optimization or a strong rigid motion assumption, we show for the first time that scene flow can also be estimated by dense interpolation of sparse matches. To this end, we find sparse…

计算机视觉与模式识别 · 计算机科学 2017-10-30 René Schuster , Oliver Wasenmüller , Georg Kuschk , Christian Bailer , Didier Stricker

The materials of surfaces in a room play an important room in shaping the auditory experience within them. Different materials absorb energy at different levels. The level of absorption also varies across frequencies. This paper…

音频与语音处理 · 电气工程与系统科学 2019-10-29 Constantinos Papayiannis , Christine Evers , Patrick A. Naylor

In this paper, we make the first attempt to align diffusion models for image inpainting with human aesthetic standards via a reinforcement learning framework, significantly improving the quality and visual appeal of inpainted images.…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Kendong Liu , Zhiyu Zhu , Chuanhao Li , Hui Liu , Huanqiang Zeng , Junhui Hou

Diffusion models have been shown to be capable of generating high-quality images, suggesting that they could contain meaningful internal representations. Unfortunately, the feature maps that encode a diffusion model's internal information…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Grace Luo , Lisa Dunlap , Dong Huk Park , Aleksander Holynski , Trevor Darrell

Diffusion models have gained attention for their ability to represent complex distributions and incorporate uncertainty, making them ideal for robust predictions in the presence of noisy or incomplete data. In this study, we develop and…

机器学习 · 计算机科学 2024-11-05 Yilin Zhuang , Sibo Cheng , Karthik Duraisamy

Methods are proposed for modifying the reverberation characteristics of sound fields in rooms by employing a loudspeaker with adjustable directivity, realized with a compact spherical loudspeaker array (SLA). These methods are based on…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Hai Morgenstern , Boaz Rafaely

Video Frame Interpolation aims to recover realistic missing frames between observed frames, generating a high-frame-rate video from a low-frame-rate video. However, without additional guidance, the large motion between frames makes this…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Jingxi Chen , Brandon Y. Feng , Haoming Cai , Tianfu Wang , Levi Burner , Dehao Yuan , Cornelia Fermuller , Christopher A. Metzler , Yiannis Aloimonos

The temporal interpolation task for 4D medical imaging, plays a crucial role in clinical practice of respiratory motion modeling. Following the simplified linear-motion hypothesis, existing approaches adopt optical flow-based models to…

图像与视频处理 · 电气工程与系统科学 2025-07-08 Xin You , Runze Yang , Chuyan Zhang , Zhongliang Jiang , Jie Yang , Nassir Navab

Neural reconstruction approaches are rapidly emerging as the preferred representation for 3D scenes, but their limited editability is still posing a challenge. In this work, we propose an approach for 3D scene inpainting -- the task of…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Ashkan Mirzaei , Riccardo De Lutio , Seung Wook Kim , David Acuna , Jonathan Kelly , Sanja Fidler , Igor Gilitschenski , Zan Gojcic

Linear systems such as room acoustics and string oscillations may be modeled as the sum of mode responses, each characterized by a frequency, damping and amplitude. Here, we consider finding the mode parameters from impulse response…

音频与语音处理 · 电气工程与系统科学 2022-02-24 Orchisama Das , Jonathan S. Abel

A primary challenge when deploying speaker recognition systems in real-world applications is performance degradation caused by environmental mismatch. We propose a diffusion-based method that takes speaker embeddings extracted from a…

音频与语音处理 · 电气工程与系统科学 2025-05-23 KiHyun Nam , Jungwoo Heo , Jee-weon Jung , Gangin Park , Chaeyoung Jung , Ha-Jin Yu , Joon Son Chung

The letter investigates the utility of text-to-image inpainting models for satellite image data. Two technical challenges of injecting structural guiding signals into the generative process as well as translating the inpainted RGB pixels to…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Mikolaj Czerkawski , Christos Tachtatzis

Video editing and generation methods often rely on pre-trained image-based diffusion models. During the diffusion process, however, the reliance on rudimentary noise sampling techniques that do not preserve correlations present in…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Pascal Chang , Jingwei Tang , Markus Gross , Vinicius C. Azevedo

Speech enhancement in ad-hoc microphone arrays is often hindered by the asynchronization of the devices composing the microphone array. Asynchronization comes from sampling time offset and sampling rate offset which inevitably occur when…

音频与语音处理 · 电气工程与系统科学 2023-08-01 Nicolas Furnon , Romain Serizel , Slim Essid , Irina Illina

A significant research effort is focused on exploiting the amazing capacities of pretrained diffusion models for the editing of images.They either finetune the model, or invert the image in the latent space of the pretrained model. However,…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Senmao Li , Joost van de Weijer , Taihang Hu , Fahad Shahbaz Khan , Qibin Hou , Yaxing Wang , Jian Yang , Ming-Ming Cheng

We present an open-access dataset of over 8000 acoustic impulse from 160 microphones spread across the body and affixed to wearable accessories. The data can be used to evaluate audio capture and array processing systems using wearable…

音频与语音处理 · 电气工程与系统科学 2019-12-12 Ryan M. Corey , Naoki Tsuda , Andrew C. Singer

Many multi-microphone speech enhancement algorithms require the relative transfer function (RTF) vector of the desired speech source, relating the acoustic transfer functions of all array microphones to a reference microphone. In this…

音频与语音处理 · 电气工程与系统科学 2022-11-22 N. Gößling , S. Doclo

Acoustic beamforming with a microphone array represents an adequate technology for remote acoustic surveillance, as the system has no mechanical parts and it has moderate size. However, in order to accomplish real implementation, several…

声音 · 计算机科学 2017-03-08 Abdulla AlShehhi , M. Luai Hammadih , M. Sami Zitouni , Saif AlKindi , Nazar Ali , Luis Weruaga