English
Related papers

Related papers: TextToon: Real-Time Text Toonify Head Avatar from …

200 papers

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li

Generating animatable and editable 3D head avatars is essential for various applications in computer vision and graphics. Traditional 3D-aware generative adversarial networks (GANs), often using implicit fields like Neural Radiance Fields…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Guohao Li , Hongyu Yang , Yifang Men , Di Huang , Weixin Li , Ruijie Yang , Yunhong Wang

In this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach employs a two-stage training…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Yanwen Wang , Yiyu Zhuang , Jiawei Zhang , Li Wang , Yifei Zeng , Xun Cao , Xinxin Zuo , Hao Zhu

We study the problem of creating a character model that can be controlled in real time from a single image of an anime character. A solution to this problem would greatly reduce the cost of creating avatars, computer games, and other…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Pramook Khungurn

Photorealistic and controllable human avatars have gained popularity in the research community thanks to rapid advances in neural rendering, providing fast and realistic synthesis tools. However, a limitation of current solutions is the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Mohamed Ilyes Lakhal , Richard Bowden

Creating relightable and animatable avatars from multi-view or monocular videos is a challenging task for digital human creation and virtual reality applications. Previous methods rely on neural radiance fields or ray tracing, resulting in…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Youyi Zhan , Tianjia Shao , He Wang , Yin Yang , Kun Zhou

Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation. However, current models do not yet convey the feeling of truly interactive communication, often generating one-way…

Machine Learning · Computer Science 2026-01-05 Taekyung Ki , Sangwon Jang , Jaehyeong Jo , Jaehong Yoon , Sung Ju Hwang

Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text instructions,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ruikui Wang , Jinheng Feng , Lang Tian , Huaishao Luo , Chaochao Li , Liangbo Zhou , Huan Zhang , Youzheng Wu , Xiaodong He

We present a novel framework for generating high-quality, animatable 4D avatar from a single image. While recent advances have shown promising results in 4D avatar creation, existing methods either require extensive multiview data or…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Fei Yin , Mallikarjun B R , Chun-Han Yao , Rafał Mantiuk , Varun Jampani

Head avatar reenactment focuses on creating animatable personal avatars from monocular videos, serving as a foundational element for applications like social signal understanding, gaming, human-machine interaction, and computer vision.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Wei Liang , Hui Yu , Derui Ding , Rachael E. Jack , Philippe G. Schyns

Recently, text-guided digital portrait editing has attracted more and more attentions. However, existing methods still struggle to maintain consistency across time, expression, and view or require specific data prerequisites. To solve these…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Haiyao Xiao , Chenglai Zhong , Xuan Gao , Yudong Guo , Juyong Zhang

While haircut indicates distinct personality, existing avatar generation methods fail to model practical hair due to the data limitation or entangled representation. We propose StrandHead, a novel text-driven method capable of generating 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Xiaokun Sun , Zeyu Cai , Ying Tai , Jian Yang , Zhenyu Zhang

We introduce FlexAvatar, a method for creating high-quality and complete 3D head avatars from a single image. A core challenge lies in the limited availability of multi-view data and the tendency of monocular training to yield incomplete 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Tobias Kirschstein , Simon Giebenhain , Matthias Nießner

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xuangeng Chu , Yuan Gan , Ziteng Cui , Shuhong Liu , Jian Wang , Bing Zhou , Tatsuya Harada

Customized text-to-video generation aims to generate text-guided videos with user-given subjects, which has gained increasing attention. However, existing works are primarily limited to single-subject oriented text-to-video generation,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Hong Chen , Xin Wang , Guanning Zeng , Yipeng Zhang , Yuwei Zhou , Feilin Han , Yaofei Wu , Wenwu Zhu

To integrate digital humans into everyday life, there is a strong demand for generating high-quality, fine-grained disentangled 3D avatars that support expressive animation and simulation capabilities, ideally from low-cost textual inputs.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Xiaokun Sun , Zhenyu Zhang , Ying Tai , Hao Tang , Zili Yi , Jian Yang

We introduce an approach that creates animatable human avatars from monocular videos using 3D Gaussian Splatting (3DGS). Existing methods based on neural radiance fields (NeRFs) achieve high-quality novel-view/novel-pose image synthesis but…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Zhiyin Qian , Shaofei Wang , Marko Mihajlovic , Andreas Geiger , Siyu Tang

The ability to animate photo-realistic head avatars reconstructed from monocular portrait video sequences represents a crucial step in bridging the gap between the virtual and real worlds. Recent advancements in head avatar techniques,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Yufan Chen , Lizhen Wang , Qijing Li , Hongjiang Xiao , Shengping Zhang , Hongxun Yao , Yebin Liu

Multimodal-driven talking face generation refers to animating a portrait with the given pose, expression, and gaze transferred from the driving image and video, or estimated from the text and audio. However, existing methods ignore the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Chao Xu , Shaoting Zhu , Junwei Zhu , Tianxin Huang , Jiangning Zhang , Ying Tai , Yong Liu

We propose HeadOn, the first real-time source-to-target reenactment approach for complete human portrait videos that enables transfer of torso and head motion, face expression, and eye gaze. Given a short RGB-D video of the target actor, we…

Computer Vision and Pattern Recognition · Computer Science 2018-05-31 Justus Thies , Michael Zollhöfer , Christian Theobalt , Marc Stamminger , Matthias Nießner