English
Related papers

Related papers: StyleLipSync: Style-based Personalized Lip-sync Vi…

200 papers

Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ziqiao Peng , Wentao Hu , Junyuan Ma , Xiangyu Zhu , Xiaomei Zhang , Hao Zhao , Hui Tian , Jun He , Hongyan Liu , Zhaoxin Fan

Text-guided image generation aimed to generate desired images conditioned on given texts, while text-guided image manipulation refers to semantically edit parts of a given image based on specified texts. For these two similar tasks, the key…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Xiaozhou You , Jian Zhang

We introduce MyStyle, a personalized deep generative prior trained with a few shots of an individual. MyStyle allows to reconstruct, enhance and edit images of a specific person, such that the output is faithful to the person's key facial…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Yotam Nitzan , Kfir Aberman , Qiurui He , Orly Liba , Michal Yarom , Yossi Gandelsman , Inbar Mosseri , Yael Pritch , Daniel Cohen-or

Taking inspiration from recent developments in visual generative tasks using diffusion models, we propose a method for end-to-end speech-driven video editing using a denoising diffusion model. Given a video of a talking person, and a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Dan Bigioi , Shubhajit Basak , Michał Stypułkowski , Maciej Zięba , Hugh Jordan , Rachel McDonnell , Peter Corcoran

The existing methods for audio-driven talking head video editing have the limitations of poor visual effects. This paper tries to tackle this problem through editing talking face images seamless with different emotions based on two modules:…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Jiacheng Su , Kunhong Liu , Liyan Chen , Junfeng Yao , Qingsong Liu , Dongdong Lv

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Xu He , Haoxian Zhang , Hejia Chen , Changyuan Zheng , Liyang Chen , Songlin Tang , Jiehui Huang , Xiaoqiang Liu , Pengfei Wan , Zhiyong Wu

This paper proposes a talking face generation method named "CP-EB" that takes an audio signal as input and a person image as reference, to synthesize a photo-realistic people talking video with head poses controlled by a short video clip…

Computer Vision and Pattern Recognition · Computer Science 2023-11-16 Jianzong Wang , Yimin Deng , Ziqi Liang , Xulong Zhang , Ning Cheng , Jing Xiao

While previous speech-driven talking face generation methods have made significant progress in improving the visual quality and lip-sync quality of the synthesized videos, they pay less attention to lip motion jitters which greatly…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Jun Ling , Xu Tan , Liyang Chen , Runnan Li , Yuchao Zhang , Sheng Zhao , Li Song

In this paper, we introduce a simple and novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead…

Graphics · Computer Science 2022-12-09 Zhentao Yu , Zixin Yin , Deyu Zhou , Duomin Wang , Finn Wong , Baoyuan Wang

Speech-driven facial video generation has been a complex problem due to its multi-modal aspects namely audio and video domain. The audio comprises lots of underlying features such as expression, pitch, loudness, prosody(speaking style) and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Shuchen Weng , Haojie Zheng , Zheng Chang , Si Li , Boxin Shi , Xinlong Wang

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

Multimedia · Computer Science 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

Generating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however, current methods lack artistic control of the style of image…

Computer Vision and Pattern Recognition · Computer Science 2022-03-02 Peter Schaldenbrand , Zhixuan Liu , Jean Oh

Text-to-image diffusion models have remarkably excelled in producing diverse, high-quality, and photo-realistic images. This advancement has spurred a growing interest in incorporating specific identities into generated content. Most…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Xiaoming Li , Xinyu Hou , Chen Change Loy

Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch). Thus, we extend the task to a more challenging setting:…

Sound · Computer Science 2025-07-28 Fang Kang , Yin Cao , Haoyu Chen

We present an algorithm for re-rendering a person from a single image under arbitrary poses. Existing methods often have difficulties in hallucinating occluded contents photo-realistically while preserving the identity and fine details in…

Computer Vision and Pattern Recognition · Computer Science 2021-09-14 Badour AlBahar , Jingwan Lu , Jimei Yang , Zhixin Shu , Eli Shechtman , Jia-Bin Huang

Facial expression transfer and reenactment has been an important research problem given its applications in face editing, image manipulation, and fabricated videos generation. We present a novel method for image-based facial expression…

Computer Vision and Pattern Recognition · Computer Science 2019-12-16 Chao Yang , Ser-Nam Lim

Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over facial animation such as speaking style and emotional expression,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Baiqin Wang , Xiangyu Zhu , Fan Shen , Hao Xu , Zhen Lei

Recent advances in audio-driven talking head generation have achieved impressive results in lip synchronization and emotional expression. However, they largely overlook the crucial task of facial attribute editing. This capability is…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Guanwen Feng , Zhiyuan Ma , Yunan Li , Jiahao Yang , Junwei Jing , Qiguang Miao

A major obstacle when attempting to train a machine learning system to evaluate facial clefts is the scarcity of large datasets of high-quality, ethics board-approved patient images. In response, we have built a deep learning-based cleft…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Abdullah Hayajneh , Erchin Serpedin , Mohammad Shaqfeh , Graeme Glass , Mitchell A. Stotland