English
Related papers

Related papers: SF-Speech: Straightened Flow for Zero-Shot Voice C…

200 papers

Diffusion-based generative models have emerged as highly effective methods for synthesizing high-quality samples. Recent works have focused on analyzing the convergence of their generation process with minimal assumptions, either through…

Machine Learning · Statistics 2025-08-25 Nishant Jain , Tong Zhang

We introduce SLED, an alternative approach to speech language modeling by encoding speech waveforms into sequences of continuous latent representations and modeling them autoregressively using an energy distance objective. The energy…

Computation and Language · Computer Science 2025-10-27 Zhengrui Ma , Yang Feng , Chenze Shao , Fandong Meng , Jie Zhou , Min Zhang

Recent advances in generative speech have increased the need for automatic detection of obviously failed synthetic outputs. This is particularly important in clinical settings such as AVATAR therapy, in which schizophrenia patients engage…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Jana Shokr , Minos Papadopoulos , Jeremy Cooperstock , Pavo Orepic

Stochastic differential equations (SDEs) are well suited to modelling noisy and irregularly sampled time series found in finance, physics, and machine learning. Traditional approaches require costly numerical solvers to sample between…

Machine Learning · Computer Science 2025-10-30 Naoki Kiyohara , Edward Johns , Yingzhen Li

Strong semantic representations improve the convergence and generation quality of diffusion and flow models. Existing approaches largely rely on external models, which require separate training, operate on misaligned objectives, and exhibit…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Hila Chefer , Patrick Esser , Dominik Lorenz , Dustin Podell , Vikash Raja , Vinh Tong , Antonio Torralba , Robin Rombach

Flow-matching models have enabled high-quality text-to-speech synthesis, but their iterative sampling process during inference incurs substantial computational cost. Although distillation is widely used to reduce the number of inference…

Sound · Computer Science 2026-02-11 Bin Lin , Peng Yang , Chao Yan , Xiaochen Liu , Wei Wang , Boyong Wu , Pengfei Tan , Xuerui Yang

Ordinary differential equation (ODE) based generative models have emerged as a powerful approach for producing high-quality samples in many applications. However, the ODE-based methods either suffer the discretization error of numerical…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Jingjing Wang , Dan Zhang , Joshua Luo , Yin Yang , Feng Luo

Diffusion probabilistic models generate samples by learning to reverse a noise-injection process that transforms data into noise. A key development is the reformulation of the reverse sampling process as a deterministic probability flow…

Machine Learning · Computer Science 2025-08-15 Daniel Zhengyu Huang , Jiaoyang Huang , Zhengjiang Lin

Generative models are capable to address difficult problems with non-unique solutions like bandwidth extension and gap filling, removing highly non-linear artifacts from codecs, clipping and distortion, as opposed to removing linear…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-18 Sebastian Braun

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.g., content,…

Ordinary differential equations (ODEs) are central to scientific modelling, but inferring their vector fields from noisy trajectories remains challenging. Current approaches such as symbolic regression, Gaussian process (GP) regression, and…

Machine Learning · Computer Science 2026-02-10 Maximilian Mauel , Johannes R. Hübers , David Berghaus , Patrick Seifner , Ramses J. Sanchez

In the landscape of modern machine learning, frozen pre-trained models provide stability and efficiency but often underperform on specific tasks due to mismatched data distributions. This paper introduces the Whisperer, a novel visual…

Machine Learning · Computer Science 2026-03-06 Samandar Samandarov , Nazirjon Ismoiljonov , Abdullah Sattorov , Temirlan Sabyrbayev

In voice conversion (VC) applications, diffusion and flow-matching models have exhibited exceptional speech quality and speaker similarity performances. However, they are limited by slow conversion owing to their iterative inference.…

Sound · Computer Science 2026-02-23 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Yuto Kondo

Current audio language models are predominantly text-first, either extending pre-trained text LLM backbones or relying on semantic-only audio tokens, limiting general audio modeling. This paper presents a systematic empirical study of…

Diffusion models are a new class of generative models that have shown outstanding performance in image generation literature. As a consequence, studies have attempted to apply diffusion models to other tasks, such as speech enhancement. A…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-10 Philippe Gonzalez , Zheng-Hua Tan , Jan Østergaard , Jesper Jensen , Tommy Sonne Alstrøm , Tobias May

Elucidating reaction mechanisms hinges on efficiently generating transition states (TSs), products, and complete reaction networks. Recent generative models, such as diffusion models for TS sampling and sequence-based architectures for…

Chemical Physics · Physics 2025-11-06 Ping Tuo , Jiale Chen , Ju Li

Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate…

Recently, sequence-to-sequence (seq-to-seq) models have been successfully applied in text-to-speech (TTS) to synthesize speech for single-language text. To synthesize speech for multiple languages usually requires multi-lingual speech from…

Sound · Computer Science 2022-11-18 Haitong Zhang , Yue Lin

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Riki Shimizu , Xilin Jiang , Nima Mesgarani

While recent zero-shot text-to-speech (TTS) models have significantly improved speech quality and expressiveness, mainstream systems still suffer from issues related to speech-text alignment modeling: 1) models without explicit speech-text…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-31 Ziyue Jiang , Yi Ren , Ruiqi Li , Shengpeng Ji , Boyang Zhang , Zhenhui Ye , Chen Zhang , Bai Jionghao , Xiaoda Yang , Jialong Zuo , Yu Zhang , Rui Liu , Xiang Yin , Zhou Zhao
‹ Prev 1 4 5 6 7 8 10 Next ›