English
Related papers

Related papers: Emotion-Aligned Generation in Diffusion Text to Sp…

200 papers

Masked diffusion language models generate text through iterative masked-token filling, but terminal-only rewards on final completions provide coarse credit assignment for the intermediate filling decisions that shape the generation process.…

Computation and Language · Computer Science 2026-05-20 Daisuke Oba , Hiroki Furuta , Naoaki Okazaki

In this project, we aim to build a Text-to-Speech system able to produce speech with a controllable emotional expressiveness. We propose a methodology for solving this problem in three main steps. The first is the collection of emotional…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-08 Noé Tits

Prompt design plays a crucial role in text-to-video (T2V) generation, yet user-provided prompts are often short, unstructured, and misaligned with training data, limiting the generative potential of diffusion-based T2V models. We present…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Bingjie Gao , Qianli Ma , Xiaoxue Wu , Shuai Yang , Guanzhou Lan , Haonan Zhao , Jiaxuan Chen , Qingyang Liu , Yu Qiao , Xinyuan Chen , Yaohui Wang , Li Niu

Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional TTS for seen…

Sound · Computer Science 2023-05-24 Minki Kang , Wooseok Han , Sung Ju Hwang , Eunho Yang

Visual generative AI models often encounter challenges related to text-image alignment and reasoning limitations. This paper presents a novel method for selectively enhancing the signal at critical denoising steps, optimizing image…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Paul Grimal , Hervé Le Borgne , Olivier Ferret

Large language models (LLMs) have been widely applied to emotional support conversation (ESC). However, complex multi-turn support remains challenging.This is because existing alignment schemes rely on sparse outcome-level signals, thus…

Computation and Language · Computer Science 2026-04-30 Chenghui Zou , Ning Wang , Tiesunlong Shen , Luwei Xiao , Chuan Ma , Xiangpeng Li , Rui Mao , Erik Cambria

Several recent end-to-end text-to-speech (TTS) models enabling single-stage training and parallel sampling have been proposed, but their sample quality does not match that of two-stage TTS systems. In this work, we present a parallel…

Sound · Computer Science 2021-06-14 Jaehyeon Kim , Jungil Kong , Juhee Son

Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptimal quality. While post-training alignment has proven…

Sound · Computer Science 2025-11-18 Junan Zhang , Xueyao Zhang , Jing Yang , Yuancheng Wang , Fan Fan , Zhizheng Wu

Diffusion models have made substantial advances in image generation, yet models trained on large, unfiltered datasets often yield outputs misaligned with human preferences. Numerous methods have been proposed to fine-tune pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Fu-Yun Wang , Yunhao Shui , Jingtan Piao , Keqiang Sun , Hongsheng Li

Large language models have revolutionized sign language generation by automatically transforming text into high-quality sign language videos, providing accessible communication for the Deaf community. However, existing LLM-based approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yanchao Zhao , Jihao Zhu , Yu Liu , Weizhuo Chen , Yuling Yang , Kun Peng

Direct Preference Optimization (DPO) has emerged as a predominant alignment method for diffusion models, facilitating off-policy training without explicit reward modeling. However, its reliance on large-scale, high-quality human preference…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Khiem Pham , Quang Nguyen , Tung Nguyen , Jingsen Zhu , Michele Santacatterina , Dimitris Metaxas , Ramin Zabih

Conditional natural language generation methods often require either expensive fine-tuning or training a large language model from scratch. Both are unlikely to lead to good results without a substantial amount of data and computational…

Computation and Language · Computer Science 2023-08-10 Yarik Menchaca Resendiz , Roman Klinger

Diffusion models have demonstrated remarkable performance in speech synthesis, but typically require multi-step sampling, resulting in low inference efficiency. Recent studies address this issue by distilling diffusion models into…

Sound · Computer Science 2025-10-08 Tao Zhu , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

Text-to-image generative models, specifically those based on diffusion models like Imagen and Stable Diffusion, have made substantial advancements. Recently, there has been a surge of interest in the delicate refinement of text prompts.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Wenyi Mo , Tianyu Zhang , Yalong Bai , Bing Su , Ji-Rong Wen , Qing Yang

Preference optimization is a standard approach to fine-tuning large language models to align with human preferences. The quantity, diversity, and representativeness of the preference dataset are critical to the effectiveness of preference…

Computation and Language · Computer Science 2025-09-18 Yuu Jinnai , Ukyo Honda

Recent work reports gains in neural text-to-speech (TTS) with Group Relative Policy Optimization (GRPO). However, in the absence of a verifiable reward for \textit{prosody}, GRPO trained on transcription-oriented signals (CER/NLL) lowers…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Seungyoun Shin , Dongha Ahn , Jiwoo Kim , Sungwook Jeon

Motion generation is essential for animating virtual characters and embodied agents. While recent text-driven methods have made significant strides, they often struggle with achieving precise alignment between linguistic descriptions and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Zhiting Gao , Dan Song , Diqiong Jiang , Chao Xue , An-An Liu

The advent of large language models (LLMs) such as ChatGPT has attracted considerable attention in various domains due to their remarkable performance and versatility. As the use of these models continues to grow, the importance of…

Neural and Evolutionary Computing · Computer Science 2024-01-19 Jill Baumann , Oliver Kramer

Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Weipeng Tan , Chuming Lin , Chengming Xu , FeiFan Xu , Xiaobin Hu , Xiaozhong Ji , Junwei Zhu , Chengjie Wang , Yanwei Fu

Text-based speech editing allows users to edit speech by intuitively cutting, copying, and pasting text to speed up the process of editing speech. In the previous work, CampNet (context-aware mask prediction network) is proposed to realize…

Sound · Computer Science 2022-12-21 Tao Wang , Jiangyan Yi , Ruibo Fu , Jianhua Tao , Zhengqi Wen , Chu Yuan Zhang