English
Related papers

Related papers: TextToon: Real-Time Text Toonify Head Avatar from …

200 papers

In this paper, we explore a reconstruction and reenactment separated framework for 3D Gaussians head, which requires only a single portrait image as input to generate controllable avatar. Specifically, we developed a large-scale one-shot…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Zhiling Ye , Cong Zhou , Xiubao Zhang , Haifeng Shen , Weihong Deng , Quan Lu

In this study, our goal is to create interactive avatar agents that can autonomously plan and animate nuanced facial movements realistically, from both visual and behavioral perspectives. Given high-level inputs about the environment and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Duomin Wang , Bin Dai , Yu Deng , Baoyuan Wang

Modeling animatable human avatars from RGB videos is a long-standing and challenging problem. Recent works usually adopt MLP-based neural radiance fields (NeRF) to represent 3D humans, but it remains difficult for pure MLPs to regress…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Zhe Li , Yipengjing Sun , Zerong Zheng , Lizhen Wang , Shengping Zhang , Yebin Liu

Photorealistic avatars have become essential for immersive applications in virtual reality (VR) and augmented reality (AR), enabling lifelike interactions in areas such as training simulations, telemedicine, and virtual collaboration. These…

Graphics · Computer Science 2025-04-18 Rendong Zhang , Alexandra Watkins , Nilanjan Sarkar

Gaussian splatting has emerged as a powerful 3D representation that harnesses the advantages of both explicit (mesh) and implicit (NeRF) 3D representations. In this paper, we seek to leverage Gaussian splatting to generate realistic…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Ye Yuan , Xueting Li , Yangyi Huang , Shalini De Mello , Koki Nagano , Jan Kautz , Umar Iqbal

Recent studies have combined 3D Gaussian and 3D Morphable Models (3DMM) to construct high-quality 3D head avatars. In this line of research, existing methods either fail to capture the dynamic textures or incur significant overhead in terms…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Yating Wang , Xuan Wang , Ran Yi , Yanbo Fan , Jichen Hu , Jingcheng Zhu , Lizhuang Ma

Diffusion-based models have gained wide adoption in the virtual human generation due to their outstanding expressiveness. However, their substantial computational requirements have constrained their deployment in real-time interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Haojie Yu , Zhaonian Wang , Yihan Pan , Meng Cheng , Hao Yang , Chao Wang , Tao Xie , Xiaoming Xu , Xiaoming Wei , Xunliang Cai

Personalized 3D avatar editing holds significant promise due to its user-friendliness and availability to applications such as AR/VR and virtual try-ons. Previous studies have explored the feasibility of 3D editing, but often struggle to…

Graphics · Computer Science 2025-04-30 Hanxi Liu , Yifang Men , Zhouhui Lian

Realistic animatable human avatars from monocular videos are crucial for advancing human-robot interaction and enhancing immersive virtual experiences. While recent research on 3DGS-based human avatars has made progress, it still struggles…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Guangan Jiang , Tianzi Zhang , Dong Li , Zhenjun Zhao , Haoang Li , Mingrui Li , Hongyu Wang

Creating photorealistic 3D head avatars from limited input has become increasingly important for applications in virtual reality, telepresence, and digital entertainment. While recent advances like neural rendering and 3D Gaussian splatting…

Graphics · Computer Science 2026-03-12 Chen Guo , Zhuo Su , Liao Wang , Jian Wang , Shuang Li , Xu Chang , Zhaohu Li , Yang Zhao , Guidong Wang , Yebin Liu , Ruqi Huang

In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextual sentiments as well as speech rhythm and pauses. To be…

Computer Vision and Pattern Recognition · Computer Science 2021-05-10 Lincheng Li , Suzhen Wang , Zhimeng Zhang , Yu Ding , Yixing Zheng , Xin Yu , Changjie Fan

Audio-driven facial animation presents an effective solution for animating digital avatars. In this paper, we detail the technical aspects of NVIDIA Audio2Face-3D, including data acquisition, network architecture, retargeting methodology,…

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Yujin Jeong , Yunji Kim , Sanghyuk Chun , Jiyoung Lee

Emerging Metaverse applications demand accessible, accurate, and easy-to-use tools for 3D digital human creations in order to depict different cultures and societies as if in the physical world. Recent large-scale vision-language advances…

Graphics · Computer Science 2023-04-07 Longwen Zhang , Qiwei Qiu , Hongyang Lin , Qixuan Zhang , Cheng Shi , Wei Yang , Ye Shi , Sibei Yang , Lan Xu , Jingyi Yu

By equipping the most recent 3D Gaussian Splatting representation with head 3D morphable models (3DMM), existing methods manage to create head avatars with high fidelity. However, most existing methods only reconstruct a head without the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Tianhao Wu , Jing Yang , Zhilin Guo , Jingyi Wan , Fangcheng Zhong , Cengiz Oztireli

Codec Avatars are a recent class of learned, photorealistic face models that accurately represent the geometry and texture of a person in 3D (i.e., for virtual reality), and are almost indistinguishable from video. In this paper we describe…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Alexander Richard , Colin Lea , Shugao Ma , Juergen Gall , Fernando de la Torre , Yaser Sheikh

Advancements in language foundation models have primarily fueled the recent surge in artificial intelligence. In contrast, generative learning of non-textual modalities, especially videos, significantly trails behind language modeling. This…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Lijun Yu

To address the ill-posed problem caused by partial observations in monocular human volumetric capture, we present AvatarCap, a novel framework that introduces animatable avatars into the capture pipeline for high-fidelity reconstruction in…

Computer Vision and Pattern Recognition · Computer Science 2022-07-13 Zhe Li , Zerong Zheng , Hongwen Zhang , Chaonan Ji , Yebin Liu

Reconstructing high-fidelity and animatable 3D head avatars from monocular videos remains a challenging yet essential task. Existing methods based on 3D Gaussian Splatting typically bind Gaussians to mesh triangles and model deformations…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Jiankuo Zhao , Xiangyu Zhu , Zidu Wang , Zhen Lei

Creating high-fidelity 3D head avatars has always been a research hotspot, but it remains a great challenge under lightweight sparse view setups. In this paper, we propose HHAvatar represented by controllable 3D Gaussians for high-fidelity…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Zhanfeng Liao , Yuelang Xu , Zhe Li , Qijing Li , Boyao Zhou , Ruifeng Bai , Di Xu , Hongwen Zhang , Yebin Liu