English
Related papers

Related papers: Zero-1-to-A: Zero-Shot One Image to Animatable Hea…

200 papers

The ability to generate diverse 3D articulated head avatars is vital to a plethora of applications, including augmented reality, cinematography, and education. Recent work on text-guided 3D object generation has shown great promise in…

Computer Vision and Pattern Recognition · Computer Science 2023-07-12 Alexander W. Bergman , Wang Yifan , Gordon Wetzstein

Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Tianyun Zhong , Chao Liang , Jianwen Jiang , Gaojie Lin , Jiaqi Yang , Zhou Zhao

Multi-view or 4D video generation has emerged as a significant research topic. Nonetheless, recent approaches to 4D generation still struggle with fundamental limitations, as they primarily rely on harnessing multiple video diffusion models…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Jangho Park , Taesung Kwon , Jong Chul Ye

Traditional methods for constructing high-quality, personalized head avatars from monocular videos demand extensive face captures and training time, posing a significant challenge for scalability. This paper introduces a novel approach to…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Zhixuan Yu , Ziqian Bai , Abhimitra Meka , Feitong Tan , Qiangeng Xu , Rohit Pandey , Sean Fanello , Hyun Soo Park , Yinda Zhang

Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end video diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Shuyuan Tu , Yueming Pan , Yinming Huang , Xintong Han , Zhen Xing , Qi Dai , Chong Luo , Zuxuan Wu , Yu-Gang Jiang

Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligned reference-pose pairs with matched skeletal structures. Handling reference-pose…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Shijun Shi , Jing Xu , Zhihang Li , Chunli Peng , Xiaoda Yang , Lijing Lu , Kai Hu , Jiangning Zhang

While high fidelity and efficiency are central to the creation of digital head avatars, recent methods relying on 2D or 3D generative models often experience limitations such as shape distortion, expression inaccuracy, and identity…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Xiaochen Zhao , Jingxiang Sun , Lizhen Wang , Jinli Suo , Yebin Liu

Portrait animation has witnessed tremendous quality improvements thanks to recent advances in video diffusion models. However, these 2D methods often compromise 3D consistency and speed, limiting their applicability in real-world scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Kaiwen Jiang , Xueting Li , Seonwook Park , Ravi Ramamoorthi , Shalini De Mello , Koki Nagano

In this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach employs a two-stage training…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Yanwen Wang , Yiyu Zhuang , Jiawei Zhang , Li Wang , Yifei Zeng , Xun Cao , Xinxin Zuo , Hao Zhu

Two major approaches exist for creating animatable human avatars. The first, a 3D-based approach, optimizes a NeRF- or 3DGS-based avatar from videos of a single person, achieving personalization through a disentangled identity…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Geonhee Sim , Gyeongsik Moon

We present a method that reconstructs and animates a 3D head avatar from a single-view portrait image. Existing methods either involve time-consuming optimization for a specific person with multiple images, or they struggle to synthesize…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Xueting Li , Shalini De Mello , Sifei Liu , Koki Nagano , Umar Iqbal , Jan Kautz

Recent large-scale text-to-image generation models have made significant improvements in the quality, realism, and diversity of the synthesized images and enable users to control the created content through language. However, the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Samaneh Azadi , Thomas Hayes , Akbar Shah , Guan Pang , Devi Parikh , Sonal Gupta

Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation. However, current models do not yet convey the feeling of truly interactive communication, often generating one-way…

Machine Learning · Computer Science 2026-01-05 Taekyung Ki , Sangwon Jang , Jaehyeong Jo , Jaehong Yoon , Sung Ju Hwang

We introduce Animate124 (Animate-one-image-to-4D), the first work to animate a single in-the-wild image into 3D video through textual motion descriptions, an underexplored problem with significant applications. Our 4D generation leverages…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Yuyang Zhao , Zhiwen Yan , Enze Xie , Lanqing Hong , Zhenguo Li , Gim Hee Lee

Building realistic and animatable avatars still requires minutes of multi-view or monocular self-rotating videos, and most methods lack precise control over gestures and expressions. To push this boundary, we address the challenge of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Jun Xiang , Yudong Guo , Leipeng Hu , Boyang Guo , Yancheng Yuan , Juyong Zhang

Diffusion-based zero-shot image restoration and enhancement models have achieved great success in various tasks of image restoration and enhancement. However, directly applying them to video restoration and enhancement results in severe…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Cong Cao , Huanjing Yue , Xin Liu , Jingyu Yang

The availability of large-scale multimodal datasets and advancements in diffusion models have significantly accelerated progress in 4D content generation. Most prior approaches rely on multiple image or video diffusion models, utilizing…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Hanwen Liang , Yuyang Yin , Dejia Xu , Hanxue Liang , Zhangyang Wang , Konstantinos N. Plataniotis , Yao Zhao , Yunchao Wei

Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistants. However, most existing approaches are limited to offline processing of complete audio…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Bohong Chen , Yumeng Li , Yinglin Xu , Youyi Zheng , Yanlin Weng , Kun Zhou

We introduce AvatarBooth, a novel method for generating high-quality 3D avatars using text prompts or specific images. Unlike previous approaches that can only synthesize avatars based on simple text descriptions, our method enables the…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Yifei Zeng , Yuanxun Lu , Xinya Ji , Yao Yao , Hao Zhu , Xun Cao

Latent diffusion models have made great strides in generating expressive portrait videos with accurate lip-sync and natural motion from a single reference image and audio input. However, these models are far from real-time, often requiring…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Hanzhong Guo , Hongwei Yi , Daquan Zhou , Alexander William Bergman , Michael Lingelbach , Yizhou Yu