English
Related papers

Related papers: TalkCLIP: Talking Head Generation with Text-Guided…

200 papers

Generative models have advanced rapidly, enabling impressive talking head generation that brings AI to life. However, most existing methods focus solely on one-way portrait animation. Even the few that support bidirectional conversational…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-25 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Zhongjian Wang , Peng Zhang , Jinwei Qi , Guangyuan Wang , Chaonan Ji , Sheng Xu , Bang Zhang , Liefeng Bo

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those methods struggle to learn…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Suzhen Wang , Lincheng Li , Yu Ding , Xin Yu

Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing emotion-aware methods rely on a single modality (either…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Phyo Thet Yee , Dimitrios Kollias , Sudeepta Mishra , Abhinav Dhall

Audio-driven emotional 3D facial animation aims to generate synchronized lip movements and vivid facial expressions. However, most existing approaches focus on static and predefined emotion labels, limiting their diversity and naturalness.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Chang Liu , Ye Pan , Chenyang Ding , Susanto Rahardja , Xiaokang Yang

Listener head generation centers on generating non-verbal behaviors (e.g., smile) of a listener in reference to the information delivered by a speaker. A significant challenge when generating such responses is the non-deterministic nature…

Graphics · Computer Science 2023-10-10 Luchuan Song , Guojun Yin , Zhenchao Jin , Xiaoyi Dong , Chenliang Xu

Owing to the power of vision-language foundation models, e.g., CLIP, the area of image synthesis has seen recent important advances. Particularly, for style transfer, CLIP enables transferring more general and abstract styles without…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Zipeng Xu , Songlong Xing , Enver Sangineto , Nicu Sebe

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and…

Sound · Computer Science 2024-12-12 Yifan Xie , Tao Feng , Xin Zhang , Xiangyang Luo , Zixuan Guo , Weijiang Yu , Heng Chang , Fei Ma , Fei Richard Yu

Hair editing is an interesting and challenging problem in computer vision and graphics. Many existing methods require well-drawn sketches or masks as conditional inputs for editing, however these interactions are neither straightforward nor…

Computer Vision and Pattern Recognition · Computer Science 2022-03-03 Tianyi Wei , Dongdong Chen , Wenbo Zhou , Jing Liao , Zhentao Tan , Lu Yuan , Weiming Zhang , Nenghai Yu

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Chenxu Zhang , Yifan Zhao , Yifei Huang , Ming Zeng , Saifeng Ni , Madhukar Budagavi , Xiaohu Guo

3D avatar creation plays a crucial role in the digital age. However, the whole production process is prohibitively time-consuming and labor-intensive. To democratize this technology to a larger audience, we propose AvatarCLIP, a zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2022-05-18 Fangzhou Hong , Mingyuan Zhang , Liang Pan , Zhongang Cai , Lei Yang , Ziwei Liu

The recently proposed visually grounded speech model SpeechCLIP is an innovative framework that bridges speech and text through images via CLIP without relying on text transcription. On this basis, this paper introduces two extensions to…

Computation and Language · Computer Science 2024-02-13 Hsuan-Fu Wang , Yi-Jen Shih , Heng-Jui Chang , Layne Berry , Puyuan Peng , Hung-yi Lee , Hsin-Min Wang , David Harwath

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazım Kemal Ekenel , Alexander Waibel

Speech-driven 3D face animation technique, extending its applications to various multimedia fields. Previous research has generated promising realistic lip movements and facial expressions from audio signals. However, traditional regression…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Ziqiao Peng , Yihao Luo , Yue Shi , Hao Xu , Xiangyu Zhu , Jun He , Hongyan Liu , Zhaoxin Fan

In this paper, we present StyleLipSync, a style-based personalized lip-sync video generative model that can generate identity-agnostic lip-synchronizing video from arbitrary audio. To generate a video of arbitrary identities, we leverage…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Taekyung Ki , Dongchan Min

Generating talking head videos through a face image and a piece of speech audio still contains many challenges. ie, unnatural head movement, distorted expression, and identity modification. We argue that these issues are mainly because of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Wenxuan Zhang , Xiaodong Cun , Xuan Wang , Yong Zhang , Xi Shen , Yu Guo , Ying Shan , Fei Wang

This paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Yifeng Ma , Jinwei Qi , Chaonan Ji , Peng Zhang , Bang Zhang , Zhidong Deng , Liefeng Bo

Recently, talking face generation has drawn ever-increasing attention from the research community in computer vision due to its arduous challenges and widespread application scenarios, e.g. movie animation and virtual anchor. Although…

Multimedia · Computer Science 2023-05-24 Jingning Xu , Benlai Tang , Mingjie Wang , Minghao Li , Meirong Ma

Virtual humans have gained considerable attention in numerous industries, e.g., entertainment and e-commerce. As a core technology, synthesizing photorealistic face frames from target speech and facial identity has been actively studied…

‹ Prev 1 8 9 10 Next ›