English
Related papers

Related papers: VectorTalker: SVG Talking Face Generation with Pro…

200 papers

We propose a two-stage framework for audio-driven talking head generation with fine-grained expression control via facial Action Units (AUs). Unlike prior methods relying on emotion labels or implicit AU conditioning, our model explicitly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Shao-Yu Chang , Jingyi Xu , Hieu Le , Dimitris Samaras

We present GStalker, a 3D audio-driven talking face generation model with Gaussian Splatting for both fast training (40 minutes) and real-time rendering (125 FPS) with a 3$\sim$5 minute video for training material, in comparison with…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Bo Chen , Shoukang Hu , Qi Chen , Chenpeng Du , Ran Yi , Yanmin Qian , Xie Chen

Modern avatar generators allow anyone to synthesize photorealistic real-time talking avatars, ushering in a new era of avatar-based human communication, such as with immersive AR/VR interactions or videoconferencing with limited bandwidths.…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Ekta Prashnani , Koki Nagano , Shalini De Mello , David Luebke , Orazio Gallo

This paper discusses the task of face-based speech synthesis, a kind of personalized speech synthesis where the synthesized voices are constrained to perceptually match with a reference face image. Due to the lack of TTS-quality…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-07 Yao Shi , Yunfei Xu , Hongbin Suo , Yulong Wan , Haifeng Liu

Speech-driven Talking Human (TH) generation, commonly known as "Talker," currently faces limitations in multi-subject driving capabilities. Extending this paradigm to "Multi-Talker," capable of animating multiple subjects simultaneously,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Yingjie Zhou , Xilei Zhu , Siyu Ren , Ziyi Zhao , Ziwen Wang , Farong Wen , Yu Zhou , Jiezhang Cao , Xiongkuo Min , Fengjiao Chen , Xiaoyu Li , Xuezhi Cao , Guangtao Zhai , Xiaohong Liu

Creating realistic, natural, and lip-readable talking face videos remains a formidable challenge. Previous research primarily concentrated on generating and aligning single-frame images while overlooking the smoothness of frame-to-frame…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Shuheng Ge , Haoyu Xing , Li Zhang , Xiangqian Wu

Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to the audio, has…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Rongliang Wu , Yingchen Yu , Fangneng Zhan , Jiahui Zhang , Xiaoqin Zhang , Shijian Lu

Current text-to-avatar methods often rely on implicit representations (e.g., NeRF, SDF, and DMTet), leading to 3D content that artists cannot easily edit and animate in graphics software. This paper introduces a novel framework for…

Graphics · Computer Science 2025-05-01 Duotun Wang , Hengyu Meng , Zeyu Cai , Zhijing Shao , Qianxi Liu , Lin Wang , Mingming Fan , Xiaohang Zhan , Zeyu Wang

Rendering photorealistic and dynamically moving human heads is crucial for ensuring a pleasant and immersive experience in AR/VR and video conferencing applications. However, existing methods often struggle to model challenging facial…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Cong Wang , Di Kang , Yan-Pei Cao , Linchao Bao , Ying Shan , Song-Hai Zhang

Recent advancements in video generation have achieved impressive motion realism, yet they often overlook character-driven storytelling, a crucial task for automated film, animation generation. We introduce Talking Characters, a more…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Cong Wei , Bo Sun , Haoyu Ma , Ji Hou , Felix Juefei-Xu , Zecheng He , Xiaoliang Dai , Luxin Zhang , Kunpeng Li , Tingbo Hou , Animesh Sinha , Peter Vajda , Wenhu Chen

Speech-driven animation has gained significant traction in recent years, with current methods achieving near-photorealistic results. However, the field remains underexplored regarding non-verbal communication despite evidence demonstrating…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Antoni Bigata Casademunt , Rodrigo Mira , Nikita Drobyshev , Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Generating talking person portraits with arbitrary speech audio is a crucial problem in the field of digital human and metaverse. A modern talking face generation method is expected to achieve the goals of generalized audio-lip…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Zhenhui Ye , Jinzheng He , Ziyue Jiang , Rongjie Huang , Jiawei Huang , Jinglin Liu , Yi Ren , Xiang Yin , Zejun Ma , Zhou Zhao

While deep learning technologies are now capable of generating realistic images confusing humans, the research efforts are turning to the synthesis of images for more concrete and application-specific purposes. Facial image generation based…

Computer Vision and Pattern Recognition · Computer Science 2020-06-11 Yeqi Bai , Tao Ma , Lipo Wang , Zhenjie Zhang

This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice. The model is based on a two-stage network. Motion cues are obtained with…

Sound · Computer Science 2022-07-20 Juan F. Montesinos , Venkatesh S. Kadandale , Gloria Haro

Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional…

Graphics · Computer Science 2026-02-27 Fangyu Du , Taiqing Li , Qian Qiao , Tan Yu , Ziwei Zhang , Dingcheng Zhen , Xu Jia , Yang Yang , Shunshun Yin , Siyuan Liu

The ability to create realistic, animatable and relightable head avatars from casual video sequences would open up wide ranging applications in communication and entertainment. Current methods either build on explicit 3D morphable meshes…

Computer Vision and Pattern Recognition · Computer Science 2023-03-01 Yufeng Zheng , Wang Yifan , Gordon Wetzstein , Michael J. Black , Otmar Hilliges

Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frame cues (e.g., emotion and head poses), while ensuring precise lip synchronization and faithful reproduction of speaking…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 He Feng , Yongjia Ma , Donglin Di , Lei Fan , Tonghua Su , Xiangqian Wu

Talking Head Generation (THG) has emerged as a transformative technology in computer vision, enabling the synthesis of realistic human faces synchronized with image, audio, text, or video inputs. This paper provides a comprehensive review…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Vineet Kumar Rakesh , Soumya Mazumdar , Research Pratim Maity , Sarbajit Pal , Amitabha Das , Tapas Samanta

Recently, generative adversarial networks have gained a lot of popularity for image generation tasks. However, such models are associated with complex learning mechanisms and demand very large relevant datasets. This work borrows concepts…

Machine Learning · Computer Science 2018-09-28 Shagan Sah , Chi Zhang , Thang Nguyen , Dheeraj Kumar Peri , Ameya Shringi , Raymond Ptucha

Speech-driven 3D talking heads generation has emerged as a significant area of interest among researchers, presenting numerous challenges. Existing methods are constrained by animating faces with fixed topologies, wherein point-wise…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Federico Nocentini , Thomas Besnier , Claudio Ferrari , Sylvain Arguillere , Stefano Berretti , Mohamed Daoudi