English
Related papers

Related papers: StableDub: Taming Diffusion Prior for Generalized …

200 papers

The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Chenxu Zhang , Chao Wang , Jianfeng Zhang , Hongyi Xu , Guoxian Song , You Xie , Linjie Luo , Yapeng Tian , Xiaohu Guo , Jiashi Feng

Dubbing is a post-production process of re-recording actors' dialogues, which is extensively used in filmmaking and video production. It is usually performed manually by professional voice actors who read lines with proper prosody, and in…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-16 Chenxu Hu , Qiao Tian , Tingle Li , Yuping Wang , Yuxuan Wang , Hang Zhao

With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still require task-specific fine-tuning and large-scale audiovisual datasets, resulting in high…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Hao Wu , Xiangyang Luo , Hao Wang , Jiawei Zhang , Yi Zhang , Jinwei Wang

Block-wise discrete diffusion offers an attractive balance between parallel generation and causal dependency modeling, making it a promising backbone for vision-language modeling. However, its practical adoption has been limited by high…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Shuang Cheng , Yuhua Jiang , Zineng Zhou , Dawei Liu , Wang Tao , Linfeng Zhang , Biqing Qi , Bowen Zhou

While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This paper proposes a new task, video-to-soundtrack (V2ST)…

Multimedia · Computer Science 2025-07-15 Wenjie Tian , Xinfa Zhu , Haohe Liu , Zhixian Zhao , Zihao Chen , Chaofan Ding , Xinhan Di , Junjie Zheng , Lei Xie

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiaxin Ye , Hongming Shan

The Stable Diffusion model is a prominent text-to-image generation model that relies on a text prompt as its input, which is encoded using the Contrastive Language-Image Pre-Training (CLIP). However, text prompts have limitations when it…

Computer Vision and Pattern Recognition · Computer Science 2024-02-16 Yuxuan Ding , Chunna Tian , Haoxuan Ding , Lingqiao Liu

Audio-visual alignment after dubbing is a challenging research problem. To this end, we propose a novel method, DubWise Multi-modal Large Language Model (LLM)-based Text-to-Speech (TTS), which can control the speech duration of synthesized…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-14 Neha Sahipjohn , Ashishkumar Gudmalwar , Nirmesh Shah , Pankaj Wasnik , Rajiv Ratn Shah

Audio-driven facial animation has made significant progress in multimedia applications, with diffusion models showing strong potential for talking-face synthesis. However, most existing works treat speech features as a monolithic…

Graphics · Computer Science 2026-04-14 Tianle Lyu , Junchuan Zhao , Ye Wang

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip. This task demands…

Sound · Computer Science 2025-03-19 Zhedong Zhang , Liang Li , Chenggang Yan , Chunshan Liu , Anton van den Hengel , Yuankai Qi

Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained by the limited word embeddings provided in pre-training,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Mingce Guo , Jingxuan He , Shengeng Tang , Zhangye Wang , Lechao Cheng

Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-quality paired audio…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Xu Wang , Shengeng Tang , Fei Wang , Lechao Cheng , Dan Guo , Feng Xue , Richang Hong

Animating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial networks (GANs), which typically suffer from notorious mode…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Lingting Zhu , Xian Liu , Xuanyu Liu , Rui Qian , Ziwei Liu , Lequan Yu

Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper, we start by…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazim Kemal Ekenel , Alexander Waibel

Diffusion models have achieved remarkable success in text-to-speech (TTS), even in zero-shot scenarios. Recent efforts aim to address the trade-off between inference speed and sound quality, often considered the primary drawback of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Changjin Han , Seokgi Lee , Gyuhyeon Nam , Gyeongsu Chae

Embodied AI agents require a fine-grained understanding of the physical world mediated through visual and language inputs. Such capabilities are difficult to learn solely from task-specific data. This has led to the emergence of pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2024-05-12 Gunshi Gupta , Karmesh Yadav , Yarin Gal , Dhruv Batra , Zsolt Kira , Cong Lu , Tim G. J. Rudner

Diffusion model based language-guided image editing has achieved great success recently. However, existing state-of-the-art diffusion models struggle with rendering correct text and text style during generation. To tackle this problem, we…

Computer Vision and Pattern Recognition · Computer Science 2023-10-19 Haoxing Chen , Zhuoer Xu , Zhangxuan Gu , Jun Lan , Xing Zheng , Yaohui Li , Changhua Meng , Huijia Zhu , Weiqiang Wang

Recent advances in diffusion-based lip-syncing generative models have demonstrated their ability to produce highly synchronized talking face videos for visual dubbing. Although these models excel at lip synchronization, they often struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yanyu Zhu , Lichen Bai , Jintao Xu , Hai-tao Zheng

Subject-driven text-to-image generation models create novel renditions of an input subject based on text prompts. Existing models suffer from lengthy fine-tuning and difficulties preserving the subject fidelity. To overcome these…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Dongxu Li , Junnan Li , Steven C. H. Hoi

While diffusion-based methods have shown impressive capabilities in capturing diverse and complex hairstyles, their ability to generate consistent and high-quality multi-view outputs -- crucial for real-world applications such as digital…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Kuiyuan Sun , Yuxuan Zhang , Jichao Zhang , Jiaming Liu , Wei Wang , Niculae Sebe , Yao Zhao