中文
相关论文

相关论文: Diff-SAGe: End-to-End Spatial Audio Generation Usi…

200 篇论文

Diffusion models have recently advanced photorealistic human synthesis, although practical talking-head generation (THG) remains constrained by high inference latency, temporal instability such as flicker and identity drift, and imperfect…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Soumya Mazumdar , Vineet Kumar Rakesh

The increasing success of audio foundation models across various tasks has led to a growing need for improved interpretability to understand their intricate decision-making processes better. Existing methods primarily focus on explaining…

声音 · 计算机科学 2024-10-11 Alican Akman , Qiyang Sun , Björn W. Schuller

We present LongCat-AudioDiT, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance. Unlike previous methods that rely on intermediate acoustic representations such as…

声音 · 计算机科学 2026-04-01 Detai Xin , Shujie Hu , Chengzuo Yang , Chen Huang , Guoqiao Yu , Guanglu Wan , Xunliang Cai

As artificial intelligence-generated content (AIGC) continues to evolve, video-to-audio (V2A) generation has emerged as a key area with promising applications in multimedia editing, augmented reality, and automated content creation. While…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Yuhuan You , Xihong Wu , Tianshu Qu

Diffusion-based generative speech enhancement (SE) has recently received attention, but reverse diffusion remains time-consuming. One solution is to initialize the reverse diffusion process with enhanced features estimated by a predictive…

In this paper we consider the problem of acoustic inversion in the context of the optoacoustic tomography image reconstruction problem. By leveraging the ability of the recently proposed diffusion models for image generative tasks among…

图像与视频处理 · 电气工程与系统科学 2024-04-17 M. G. González , M. Vera , A. Dreszman , L. J. Rey Vega

We introduce an approach to convert mono audio recorded by a 360 video camera into spatial audio, a representation of the distribution of sound over the full viewing sphere. Spatial audio is an important component of immersive 360 video…

声音 · 计算机科学 2018-09-10 Pedro Morgado , Nuno Vasconcelos , Timothy Langlois , Oliver Wang

This contribution introduces a dataset of 7th-order Ambisonic Room Impulse Responses (HOA-RIRs), created using the Image Source Method. By employing higher-order Ambisonics, our dataset enables precise spatial audio reproduction, a critical…

声音 · 计算机科学 2025-06-02 Shivam Saini , Jürgen Peissig

This paper presents DiffMoog - a differentiable modular synthesizer with a comprehensive set of modules typically found in commercial instruments. Being differentiable, it allows integration into neural networks, enabling automated sound…

音频与语音处理 · 电气工程与系统科学 2024-01-24 Noy Uzrad , Oren Barkan , Almog Elharar , Shlomi Shvartzman , Moshe Laufer , Lior Wolf , Noam Koenigstein

Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.g., mel-spectrogram) given a music score. Previous singing acoustic…

音频与语音处理 · 电气工程与系统科学 2022-03-23 Jinglin Liu , Chengxi Li , Yi Ren , Feiyang Chen , Zhou Zhao

The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video processing. First, we…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Onkar Susladkar , Jishu Sen Gupta , Chirag Sehgal , Sparsh Mittal , Rekha Singhal

Properly setting up recording conditions, including microphone type and placement, room acoustics, and ambient noise, is essential to obtaining the desired acoustic characteristics of speech. In this paper, we propose Diff-R-EN-T, a…

声音 · 计算机科学 2024-01-17 Jaekwon Im , Juhan Nam

Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically rely on synthetic data pipelines, which may not reflect…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Runwu Shi , Kai Li , Chang Li , Jiang Wang , Sihan Tan , Kazuhiro Nakadai

While diffusion models are best known for their performance in generative tasks, they have also been successfully applied to many other tasks, including audio source separation. However, current generative approaches to music source…

音频与语音处理 · 电气工程与系统科学 2026-04-24 Yun-Ning , Hung , Richard Vogl , Filip Korzeniowski , Igor Pereira

While 3D Gaussian representations (3DGS) have proven effective for modeling the geometry and appearance of objects, their potential for capturing other physical attributes-such as sound-remains largely unexplored. In this paper, we present…

声音 · 计算机科学 2025-07-29 Chunshi Wang , Hongxing Li , Yawei Luo

Although recent works on neural vocoder have improved the quality of synthesized audio, there still exists a gap between generated and ground-truth audio in frequency space. This difference leads to spectral artifacts such as hissing noise…

音频与语音处理 · 电气工程与系统科学 2021-06-15 Ji-Hoon Kim , Sang-Hoon Lee , Ji-Hyun Lee , Seong-Whan Lee

Generative models have excelled in audio tasks using approaches such as language models, diffusion, and flow matching. However, existing generative approaches for speech enhancement (SE) face notable challenges: language model-based methods…

音频与语音处理 · 电气工程与系统科学 2025-05-28 Ziqian Wang , Zikai Liu , Xinfa Zhu , Yike Zhu , Mingshuai Liu , Jun Chen , Longshuai Xiao , Chao Weng , Lei Xie

Spatial audio is fundamental to immersive virtual experiences, yet synthesizing high-fidelity binaural audio from sparse observations remains a significant challenge. Existing methods typically rely on implicit neural representations…

声音 · 计算机科学 2026-04-13 Chunhao Bi , Houqiang Zhong , Zhixin Xu , Li Song , Zhengxue Cheng

Recent progress in diffusion-based audio generation and restoration has substantially improved performance across heterogeneous conditioning regimes, including text-conditioned audio generation and audio-conditioned super-resolution.…

声音 · 计算机科学 2026-05-07 Xuanhao Zhang , Chang Li

This paper introduces an audio-visual speech enhancement system that leverages score-based generative models, also known as diffusion models, conditioned on visual information. In particular, we exploit audio-visual embeddings obtained from…

音频与语音处理 · 电气工程与系统科学 2023-06-05 Julius Richter , Simone Frintrop , Timo Gerkmann