English
Related papers

Related papers: DiffAVA: Personalized Text-to-Audio Generation wit…

200 papers

We are interested in a novel task, namely low-resource text-to-talking avatar. Given only a few-minute-long talking person video with the audio track as the training data and arbitrary texts as the driving input, we aim to synthesize…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Zhenhui Ye , Ziyue Jiang , Yi Ren , Jinglin Liu , Chen Zhang , Xiang Yin , Zejun Ma , Zhou Zhao

Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image…

Machine Learning · Computer Science 2025-03-18 Xulin Fan , Heting Gao , Ziyi Chen , Peng Chang , Mei Han , Mark Hasegawa-Johnson

Diffusion models have exhibited promising progress in video generation. However, they often struggle to retain consistent details within local regions across frames. One underlying cause is that traditional diffusion models approximate…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yupu Yao , Shangqi Deng , Zihan Cao , Harry Zhang , Liang-Jian Deng

Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is challenging as user-input prompts are often under-specified…

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

Graphics · Computer Science 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang

This paper proposes VARA-TTS, a non-autoregressive (non-AR) text-to-speech (TTS) model using a very deep Variational Autoencoder (VDVAE) with Residual Attention mechanism, which refines the textual-to-acoustic alignment layer-wisely.…

Sound · Computer Science 2021-02-15 Peng Liu , Yuewen Cao , Songxiang Liu , Na Hu , Guangzhi Li , Chao Weng , Dan Su

Recent advancements in text-to-image models, particularly diffusion models, have shown significant promise. However, compositional text-to-image models frequently encounter difficulties in generating high-quality images that accurately…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Song Wen , Guian Fang , Renrui Zhang , Peng Gao , Hao Dong , Dimitris Metaxas

Text-to-speech(TTS) has undergone remarkable improvements in performance, particularly with the advent of Denoising Diffusion Probabilistic Models (DDPMs). However, the perceived quality of audio depends not solely on its content, pitch,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-23 Huadai Liu , Rongjie Huang , Xuan Lin , Wenqiang Xu , Maozong Zheng , Hong Chen , Jinzheng He , Zhou Zhao

The Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both audio and video elements. To support research in this field,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Yuxin Mao , Xuyang Shen , Jing Zhang , Zhen Qin , Jinxing Zhou , Mochu Xiang , Yiran Zhong , Yuchao Dai

Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visual signals lack clear semantic correspondence. We propose to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Seongah Kim , Dinh Phu Tran , Hyeontaek Hwang , Saad Wazir , Duc Do Minh , Daeyoung Kim

The advancements in generative modeling, particularly the advent of diffusion models, have sparked a fundamental question: how can these models be effectively used for discriminative tasks? In this work, we find that generative models can…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Mihir Prabhudesai , Tsung-Wei Ke , Alexander C. Li , Deepak Pathak , Katerina Fragkiadaki

Diffusion models are generative models with impressive text-to-image synthesis capabilities and have spurred a new wave of creative methods for classical machine learning tasks. However, the best way to harness the perceptual knowledge of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Neehar Kondapaneni , Markus Marks , Manuel Knott , Rogerio Guimaraes , Pietro Perona

In recent years, image generation has shown a great leap in performance, where diffusion models play a central role. Although generating high-quality images, such models are mainly conditioned on textual descriptions. This begs the…

Sound · Computer Science 2023-05-23 Guy Yariv , Itai Gat , Lior Wolf , Yossi Adi , Idan Schwartz

Objective: While recent advances in text-conditioned generative models have enabled the synthesis of realistic medical images, progress has been largely confined to 2D modalities such as chest X-rays. Extending text-to-image generation to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Daniele Molino , Camillo Maria Caruso , Filippo Ruffini , Paolo Soda , Valerio Guarrasi

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. However, there are few…

Computer Vision and Pattern Recognition · Computer Science 2023-04-21 Shuai Shen , Wenliang Zhao , Zibin Meng , Wanhua Li , Zheng Zhu , Jie Zhou , Jiwen Lu

This paper presents VoiceLDM, a model designed to produce audio that accurately follows two distinct natural language text prompts: the description prompt and the content prompt. The former provides information about the overall…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-26 Yeonghyeon Lee , Inmo Yeon , Juhan Nam , Joon Son Chung

We present xGen-VideoSyn-1, a text-to-video (T2V) generation model capable of producing realistic scenes from textual descriptions. Building on recent advancements, such as OpenAI's Sora, we explore the latent diffusion model (LDM)…

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio…

Diffusion-based text-to-audio (TTA) generation has made substantial progress, leveraging latent diffusion model (LDM) to produce high-quality, diverse and instruction-relevant audios. However, beyond generation, the task of audio editing…

Sound · Computer Science 2024-10-01 Yuhang Jia , Yang Chen , Jinghua Zhao , Shiwan Zhao , Wenjia Zeng , Yong Chen , Yong Qin

Despite significant advancements in Text-to-Audio (TTA) generation models achieving high-fidelity audio with fine-grained context understanding, they struggle to model the relations between audio events described in the input text. However,…

Machine Learning · Computer Science 2026-04-10 Yuhang He , Yash Jain , Xubo Liu , Andrew Markham , Vibhav Vineet