中文
相关论文

相关论文: DiffAU: Diffusion-Based Ambisonics Upscaling

200 篇论文

A source separation method using a full-rank spatial covariance model has been proposed by Duong et al. ["Under-determined Reverberant Audio Source Separation Using a Full-rank Spatial Covariance Model," IEEE Trans. ASLP, vol. 18, no. 7,…

声音 · 计算机科学 2018-05-18 Nobutaka Ito , Shoko Araki , Tomohiro Nakatani

Most existing integrated sensing and communication (ISAC) studies focus on enabling a base station (BS) to support sensing and communication over shared resources through advanced waveform design and power allocation. In contrast, the…

信号处理 · 电气工程与系统科学 2026-05-25 Noor Waqar , Kai-Kit Wong , Chan-Byoung Chae , Ross Murch

Retrieving 3D bone anatomy from biplanar X-ray images is crucial since it can significantly reduce radiation exposure compared to traditional CT-based methods. Although various deep learning models have been proposed to address this complex…

图像与视频处理 · 电气工程与系统科学 2024-07-23 Jixiang Chen , Yiqun Lin , Haoran Sun , Xiaomeng Li

The goal of speech enhancement (SE) is to eliminate the background interference from the noisy speech signal. Generative models such as diffusion models (DM) have been applied to the task of SE because of better generalization in unseen…

声音 · 计算机科学 2023-09-06 Wen Wang , Dongchao Yang , Qichen Ye , Bowen Cao , Yuexian Zou

Multichannel speech enhancement leverages spatial cues to improve intelligibility and quality, but most learning-based methods rely on specific microphone array geometry, unable to account for geometry changes. To mitigate this limitation,…

音频与语音处理 · 电气工程与系统科学 2025-09-19 Michael Tatarjitzky , Boaz Rafaely

Deep learning-based direction-of-arrival (DoA) estimation has gained increasing popularity. A popular family of DoA estimation algorithms is beamforming methods, which operate by constructing a spatial filter that is applied to array…

计算工程、金融与科学 · 计算机科学 2025-12-25 Xuyao Deng , Yong Dou , Kele Xu

Generative diffusion processes are an emerging and effective tool for image and speech generation. In the existing methods, the underline noise distribution of the diffusion process is Gaussian noise. However, fitting distributions with…

机器学习 · 计算机科学 2021-06-17 Eliya Nachmani , Robin San Roman , Lior Wolf

The objective of this work is to extract target speaker's voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining…

音频与语音处理 · 电气工程与系统科学 2023-10-31 Suyeon Lee , Chaeyoung Jung , Youngjoon Jang , Jaehun Kim , Joon Son Chung

In the rapidly evolving fields of virtual and augmented reality, accurate spatial audio capture and reproduction are essential. For these applications, Ambisonics has emerged as a standard format. However, existing methods for encoding…

音频与语音处理 · 电气工程与系统科学 2024-11-27 Yhonatan Gayer , Vladimir Tourbabin , Zamir Ben-Hur , Jacob Donley , Boaz Rafaely

Unsupervised Anomaly Detection (UAD) techniques aim to identify and localize anomalies without relying on annotations, only leveraging a model trained on a dataset known to be free of anomalies. Diffusion models learn to modify inputs $x$…

计算机视觉与模式识别 · 计算机科学 2024-03-06 Sergio Naval Marimont , Matthew Baugh , Vasilis Siomos , Christos Tzelepis , Bernhard Kainz , Giacomo Tarroni

Spatial time series imputation is critically important to many real applications such as intelligent transportation and air quality monitoring. Although recent transformer and diffusion model based approaches have achieved significant…

机器学习 · 计算机科学 2023-09-06 Shunyang Zhang , Senzhang Wang , Xianzhen Tan , Ruochen Liu , Jian Zhang , Jianxin Wang

Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified…

声音 · 计算机科学 2023-10-03 Muqiao Yang , Chunlei Zhang , Yong Xu , Zhongweiyang Xu , Heming Wang , Bhiksha Raj , Dong Yu

A primary challenge in developing synthetic spatial hearing systems, particularly underwater, is accurately modeling sound scattering. Biological organisms achieve 3D spatial hearing by exploiting sound scattering off their bodies to…

声音 · 计算机科学 2026-03-03 Siminfar Samakoush Galougah , Pranav Pulijala , Ramani Duraiswami

Optimizing high-dimensional and complex black-box functions is crucial in numerous scientific applications. While Bayesian optimization (BO) is a powerful method for sample-efficient optimization, it struggles with the curse of…

机器学习 · 计算机科学 2025-07-08 Taeyoung Yun , Kiyoung Om , Jaewoo Lee , Sujin Yun , Jinkyoo Park

Diffusion models (DMs) have revolutionized generative learning. They utilize a diffusion process to encode data into a simple Gaussian distribution. However, encoding a complex, potentially multimodal data distribution into a single…

机器学习 · 计算机科学 2024-07-04 Yilun Xu , Gabriele Corso , Tommi Jaakkola , Arash Vahdat , Karsten Kreis

Adversarial Imitation Learning is traditionally framed as a two-player zero-sum game between a learner and an adversarially chosen cost function, and can therefore be thought of as the sequential generalization of a Generative Adversarial…

机器学习 · 计算机科学 2025-03-04 Runzhe Wu , Yiding Chen , Gokul Swamy , Kianté Brantley , Wen Sun

We introduce HiFi-HARP, a large-scale dataset of 7th-order Higher-Order Ambisonic Room Impulse Responses (HOA-RIRs) consisting of more than 100,000 RIRs generated via a hybrid acoustic simulation in realistic indoor scenes. HiFi-HARP…

声音 · 计算机科学 2025-10-27 Shivam Saini , Jürgen Peissig

Diffusion autoencoders (DAs) are variants of diffusion generative models that use an input-dependent latent variable to capture representations alongside the diffusion process. These representations, to varying extents, can be used for…

机器学习 · 计算机科学 2025-06-03 Magdalena Proszewska , Nikolay Malkin , N. Siddharth

This work introduces a feature extracted from stereophonic/binaural audio signals aiming to represent a measure of perceived quality degradation in processed spatial auditory scenes. The feature extraction technique is based on a simplified…

音频与语音处理 · 电气工程与系统科学 2022-12-06 Pablo M. Delgado , Jürgen Herre

Asymptotic multiple scale homogenisation allows to determine the effective behaviour of a porous medium by starting from the pore-scale description, when there is a large separation between the pore-scale and the macroscopic scale. When the…

经典物理 · 物理学 2019-03-06 Pascale Royer