中文
相关论文

相关论文: DiTSE: High-Fidelity Generative Speech Enhancement…

200 篇论文

Latent diffusion models have shown promising results in audio generation, making notable advancements over traditional methods. However, their performance, while impressive with short audio clips, faces challenges when extended to longer…

声音 · 计算机科学 2024-07-16 Zhenxiong Tan , Xinyin Ma , Gongfan Fang , Xinchao Wang

Recently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when generalising to…

计算机视觉与模式识别 · 计算机科学 2023-09-20 Berné Nortier , Mostafa Sadeghi , Romain Serizel

Target Speaker Extraction (TSE) aims to isolate a specific speaker's voice from a mixture, guided by a pre-recorded enrollment. While TSE bypasses the global permutation ambiguity of blind source separation, it remains vulnerable to speaker…

音频与语音处理 · 电气工程与系统科学 2026-04-10 Zikai Liu , Ziqian Wang , Xingchen Li , Yike Zhu , Shuai Wang , Longshuai Xiao , Lei Xie

Neural Text-to-Speech (TTS) systems find broad applications in voice assistants, e-learning, and audiobook creation. The pursuit of modern models, like Diffusion Models (DMs), holds promise for achieving high-fidelity, real-time speech…

声音 · 计算机科学 2024-04-02 Xiang Li , Fan Bu , Ambuj Mehrish , Yingting Li , Jiale Han , Bo Cheng , Soujanya Poria

Recently, denoising diffusion models have demonstrated remarkable performance among generative models in various domains. However, in the speech domain, the application of diffusion models for synthesizing time-varying audio faces…

音频与语音处理 · 电气工程与系统科学 2023-06-13 Ji-Sang Hwang , Sang-Hoon Lee , Seong-Whan Lee

Animating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial networks (GANs), which typically suffer from notorious mode…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Lingting Zhu , Xian Liu , Xuanyu Liu , Rui Qian , Ziwei Liu , Lequan Yu

This paper introduces a discrete diffusion model (DDM) framework for text-aligned speech tokenization and reconstruction. By replacing the auto-regressive speech decoder with a discrete diffusion counterpart, our model achieves…

音频与语音处理 · 电气工程与系统科学 2025-09-25 Pin-Jui Ku , He Huang , Jean-Marie Lemercier , Subham Sekhar Sahoo , Zhehuai Chen , Ante Jukić

Diffusion models have shown strong performance in speech enhancement, but their real-time applicability has been limited by multi-step iterative sampling. Consistency distillation has recently emerged as a promising alternative by…

音频与语音处理 · 电气工程与系统科学 2026-05-19 Liang Xu , Longfei Felix Yan , W. Bastiaan Kleijn

Inverse problems arise in a multitude of applications, where the goal is to recover a clean signal from noisy and possibly (non)linear observations. The difficulty of a reconstruction problem depends on multiple factors, such as the ground…

图像与视频处理 · 电气工程与系统科学 2024-08-21 Zalan Fabian , Berk Tinaz , Mahdi Soltanolkotabi

Diffusion large language models (dLLMs) have recently attracted significant attention for their ability to enhance diversity, controllability, and parallelism. However, their non-sequential, bidirectionally masked generation makes quality…

计算与语言 · 计算机科学 2026-03-04 Linhao Zhong , Linyu Wu , Wen Wang , Yuling Xi , Chenchen Jing , Jiaheng Zhang , Hao Chen , Chunhua Shen

In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flow-matching generators have improved target-speech fidelity.…

声音 · 计算机科学 2026-03-12 Duojia Li , Shuhan Zhang , Zihan Qian , Wenxuan Wu , Shuai Wang , Qingyang Hong , Lin Li , Haizhou Li

Benefiting from the development of deep learning, text-to-speech (TTS) techniques using clean speech have achieved significant performance improvements. The data collected from real scenes often contains noise and generally needs to be…

音频与语音处理 · 电气工程与系统科学 2023-09-06 Qiushi Zhu , Yu Gu , Rilin Chen , Chao Weng , Yuchen Hu , Lirong Dai , Jie Zhang

Scene text editing is a challenging task that involves modifying or inserting specified texts in an image while maintaining its natural and realistic appearance. Most previous approaches to this task rely on style-transfer models that crop…

计算机视觉与模式识别 · 计算机科学 2023-04-13 Jiabao Ji , Guanhua Zhang , Zhaowen Wang , Bairu Hou , Zhifei Zhang , Brian Price , Shiyu Chang

Speech enhancement (SE) aims to extract the clean waveform from noise-contaminated measurements to improve the speech quality and intelligibility. Although learning-based methods can perform much better than traditional counterparts, the…

音频与语音处理 · 电气工程与系统科学 2024-09-23 Haoyin Yan , Jie Zhang , Cunhang Fan , Yeping Zhou , Peiqi Liu

Speech enhancement techniques based on deep learning have brought significant improvement on speech quality and intelligibility. Nevertheless, a large gain in speech quality measured by objective metrics, such as perceptual evaluation of…

音频与语音处理 · 电气工程与系统科学 2020-07-06 Bo Wu , Meng Yu , Lianwu Chen , Yong Xu , Chao Weng , Dan Su , Dong Yu

There has been a growing effort to develop universal speech enhancement (SE) to handle inputs with various speech distortions and recording conditions. The URGENT Challenge series aims to foster such universal SE by embracing a broad range…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Kohei Saijo , Wangyou Zhang , Samuele Cornell , Robin Scheibler , Chenda Li , Zhaoheng Ni , Anurag Kumar , Marvin Sach , Yihui Fu , Wei Wang , Tim Fingscheidt , Shinji Watanabe

Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, and styles (e.g., singing). Current large TTS systems usually…

音频与语音处理 · 电气工程与系统科学 2023-05-31 Kai Shen , Zeqian Ju , Xu Tan , Yanqing Liu , Yichong Leng , Lei He , Tao Qin , Sheng Zhao , Jiang Bian

In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising…

声音 · 计算机科学 2025-07-18 Luca Della Libera , Cem Subakan , Mirco Ravanelli

We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to robustly separate sources and extract an initial clean speech…

声音 · 计算机科学 2025-06-03 Nabarun Goswami , Tatsuya Harada

Generating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering,…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Maksim Siniukov , Di Chang , Minh Tran , Hongkun Gong , Ashutosh Chaubey , Mohammad Soleymani