English
Related papers

Related papers: DFADD: The Diffusion and Flow-Matching Based Audio…

200 papers

Neural Text-to-Speech (TTS) systems find broad applications in voice assistants, e-learning, and audiobook creation. The pursuit of modern models, like Diffusion Models (DMs), holds promise for achieving high-fidelity, real-time speech…

Sound · Computer Science 2024-04-02 Xiang Li , Fan Bu , Ambuj Mehrish , Yingting Li , Jiale Han , Bo Cheng , Soujanya Poria

We present novel approaches involving generative adversarial networks and diffusion models in order to synthesize high quality, live and spoof fingerprint images while preserving features such as uniqueness and diversity. We generate live…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 W. Tang , D. Figueroa , D. Liu , K. Johnsson , A. Sopasakis

Diffusion probabilistic models have demonstrated an outstanding capability to model natural images and raw audio waveforms through a paired diffusion and reverse processes. The unique property of the reverse process (namely, eliminating…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-23 Yen-Ju Lu , Yu Tsao , Shinji Watanabe

Ambient diffusion is a recently proposed framework for training diffusion models using corrupted data. Both Ambient Diffusion and alternative SURE-based approaches for learning diffusion models from corrupted data resort to approximations…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Giannis Daras , Alexandros G. Dimakis , Constantinos Daskalakis

Modern deep generative models can now produce high-quality synthetic samples that are often indistinguishable from real training data. A growing body of research aims to understand why recent methods, such as diffusion and flow matching…

Machine Learning · Computer Science 2025-12-03 Quentin Bertrand , Anne Gagneux , Mathurin Massias , Rémi Emonet

Speaker-specific anti-spoofing and synthesis-source tracing are central challenges in audio anti-spoofing. Progress has been hampered by the lack of datasets that systematically vary model architectures, synthesis pipelines, and generative…

Sound · Computer Science 2026-01-14 Surya Subramani , Hashim Ali , Hafiz Malik

Humans use context to assess the veracity of information. However, current audio deepfake detectors only analyze the audio file without considering either context or transcripts. We create and analyze a Journalist-provided Deepfake Dataset…

A large and growing amount of speech content in real-life scenarios is being recorded on consumer-grade devices in uncontrolled environments, resulting in degraded speech quality. Transforming such low-quality device-degraded speech into…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-23 Haoyu Li , Junichi Yamagishi

Speaker embedding based zero-shot Text-to-Speech (TTS) systems enable high-quality speech synthesis for unseen speakers using minimal data. However, these systems are vulnerable to adversarial attacks, where an attacker introduces…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-07 Ze Li , Yao Shi , Yunfei Xu , Ming Li

The rapid growth of the digital economy in South-East Asia (SEA) has amplified the risks of audio deepfakes, yet current datasets cover SEA languages only sparsely, leaving models poorly equipped to handle this critical region. This…

Sound · Computer Science 2025-09-26 Jinyang Wu , Nana Hou , Zihan Pan , Qiquan Zhang , Sailor Hardik Bhupendra , Soumik Mondal

Flow matching as a paradigm of generative model achieves notable success across various domains. However, existing methods use either multi-round training or knowledge within minibatches, posing challenges in finding a favorable coupling…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Siyu Xing , Jie Cao , Huaibo Huang , Haichao Shi , Xiao-Yu Zhang

Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Tianyun Zhong , Chao Liang , Jianwen Jiang , Gaojie Lin , Jiaqi Yang , Zhou Zhao

Diffusion models have gained attention in speech enhancement tasks, providing an alternative to conventional discriminative methods. However, research on target speech extraction under multi-speaker noisy conditions remains relatively…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Leying Zhang , Yao Qian , Linfeng Yu , Heming Wang , Hemin Yang , Long Zhou , Shujie Liu , Yanmin Qian

Recent progress in large-scale zero-shot speech synthesis has been significantly advanced by language models and diffusion models. However, the generation process of both methods is slow and computationally intensive. Efficient speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-25 Zhen Ye , Zeqian Ju , Haohe Liu , Xu Tan , Jianyi Chen , Yiwen Lu , Peiwen Sun , Jiahao Pan , Weizhen Bian , Shulin He , Wei Xue , Qifeng Liu , Yike Guo

Traditional anti-spoofing focuses on models and datasets built on synthetic speech with mostly neutral state, neglecting diverse emotional variations. As a result, their robustness against high-quality, emotionally expressive synthetic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-02 Aurosweta Mahapatra , Ismail Rasim Ulgen , Abinay Reddy Naini , Carlos Busso , Berrak Sisman

Dataset distillation seeks to synthesize a highly compact dataset that achieves performance comparable to the original dataset on downstream tasks. For the classification task that use pre-trained self-supervised models as backbones,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Qianxin Xia , Jiawei Du , Xin Zhang , Yuhan Zhang , Jielei Wang , Guoming Lu

This paper describes the BUT submission to the ESDD 2026 Challenge, specifically focusing on Track 1: Environmental Sound Deepfake Detection with Unseen Generators. To address the critical challenge of generalizing to audio generated by…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-10 Junyi Peng , Lin Zhang , Jin Li , Oldrich Plchot , Jan Cernocky

Creating synthetic voices with found data is challenging, as real-world recordings often contain various types of audio degradation. One way to address this problem is to pre-enhance the speech with an enhancement model and then use the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-03 Yusheng Tian , Wei Liu , Tan Lee

Diffusion models have emerged as powerful deep generative techniques, producing high-quality and diverse samples in applications in various domains including audio. While existing reviews provide overviews, there remains limited in-depth…

Sound · Computer Science 2026-01-16 Ge Zhu , Yutong Wen , Zhiyao Duan

Flow-Matching (FM)-based zero-shot text-to-speech (TTS) systems exhibit high-quality speech synthesis and robust generalization capabilities. However, the speaker representation ability of such systems remains underexplored, primarily due…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-18 Haoyu Li , Mingyang Han , Yu Xi , Dongxiao Wang , Hankun Wang , Haoxiang Shi , Boyu Li , Jun Song , Bo Zheng , Shuai Wang , Kai Yu
‹ Prev 1 8 9 10 Next ›