English
Related papers

Related papers: LiftAvatar: Kinematic-Space Completion for Express…

200 papers

We present Better Together, a method that simultaneously solves the human pose estimation problem while reconstructing a photorealistic 3D human avatar from multi-view videos. While prior art usually solves these problems separately, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Arthur Moreau , Mohammed Brahimi , Richard Shaw , Athanasios Papaioannou , Thomas Tanay , Zhensong Zhang , Eduardo Pérez-Pellitero

Reconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Jiawei Zhang , Zijian Wu , Zhiyang Liang , Yicheng Gong , Dongfang Hu , Yao Yao , Xun Cao , Hao Zhu

Multi-view volumetric rendering techniques have recently shown great potential in modeling and synthesizing high-quality head avatars. A common approach to capture full head dynamic performances is to track the underlying geometry using a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Kartik Teotia , Mallikarjun B R , Xingang Pan , Hyeongwoo Kim , Pablo Garrido , Mohamed Elgharib , Christian Theobalt

Photorealistic avatars have become essential for immersive applications in virtual reality (VR) and augmented reality (AR), enabling lifelike interactions in areas such as training simulations, telemedicine, and virtual collaboration. These…

Graphics · Computer Science 2025-04-18 Rendong Zhang , Alexandra Watkins , Nilanjan Sarkar

Efficiently reconstructing 3D scenes from monocular video remains a core challenge in computer vision, vital for applications in virtual reality, robotics, and scene understanding. Recently, frame-by-frame progressive reconstruction without…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Wenyan Cong , Hanqing Zhu , Kevin Wang , Jiahui Lei , Colton Stearns , Yuanhao Cai , Leonidas Guibas , Zhangyang Wang , Zhiwen Fan

Realistic digital avatars require expressive and dynamic hair motion; however, most existing head avatar methods assume rigid hair movement. These methods often fail to disentangle hair from the head, representing it as a simple outer shell…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Berna Kabadayi , Vanessa Sklyarova , Wojciech Zielonka , Justus Thies , Gerard Pons-Moll

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

We present FaceLift, a novel feed-forward approach for generalizable high-quality 360-degree 3D head reconstruction from a single image. Our pipeline first employs a multi-view latent diffusion model to generate consistent side and back…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Weijie Lyu , Yi Zhou , Ming-Hsuan Yang , Zhixin Shu

While 2D pose estimation has advanced our ability to interpret body movements in animals and primates, it is limited by the lack of depth information, constraining its application range. 3D pose estimation provides a more comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Soumyaratna Debnath , Harish Katti , Shashikant Verma , Shanmuganathan Raman

Creating realistic avatars from a single RGB image is an attractive yet challenging problem. Due to its ill-posed nature, recent works leverage powerful prior from 2D diffusion models pretrained on large datasets. Although 2D diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yuxuan Xue , Xianghui Xie , Riccardo Marin , Gerard Pons-Moll

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have achieved notable…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yihong Lin , Zhaoxin Fan , Xianjia Wu , Lingyu Xiong , Liang Peng , Xiandong Li , Wenxiong Kang , Songju Lei , Huang Xu

Generating high-fidelity 3D avatars from text or image prompts is highly sought after in virtual reality and human-computer interaction. However, existing text-driven methods often rely on iterative Score Distillation Sampling (SDS) or CLIP…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Hong Li , Yutang Feng , Minqi Meng , Yichen Yang , Xuhui Liu , Baochang Zhang

In this paper, we introduce GaussianMotion, a novel human rendering model that generates fully animatable scenes aligned with textual descriptions using Gaussian Splatting. Although existing methods achieve reasonable text-to-3D generation…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Gyumin Shim , Sangmin Lee , Jaegul Choo

Efficiently estimating the full-body pose with minimal wearable devices presents a worthwhile research direction. Despite significant advancements in this field, most current research neglects to explore full-body avatar estimation under…

Human-Computer Interaction · Computer Science 2024-07-03 Bo Qian , Zhenhuan Wei , Jiashuo Li , Xing Wei

Existing methods for image-to-3D avatar generation struggle to produce highly detailed, animation-ready avatars suitable for real-world applications. We introduce AdaHuman, a novel framework that generates high-fidelity animatable 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Yangyi Huang , Ye Yuan , Xueting Li , Jan Kautz , Umar Iqbal

We introduce AvatarBooth, a novel method for generating high-quality 3D avatars using text prompts or specific images. Unlike previous approaches that can only synthesize avatars based on simple text descriptions, our method enables the…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Yifei Zeng , Yuanxun Lu , Xinya Ji , Yao Yao , Hao Zhu , Xun Cao

Despite recent progress in 3D Gaussian-based head avatar modeling, efficiently generating high fidelity avatars remains a challenge. Current methods typically rely on extensive multi-view capture setups or monocular videos with per-identity…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Xinya Ji , Sebastian Weiss , Manuel Kansy , Jacek Naruniec , Xun Cao , Barbara Solenthaler , Derek Bradley

In this paper, we present a novel method that facilitates the creation of vivid 3D Gaussian avatars from monocular video inputs (GVA). Our innovation lies in addressing the intricate challenges of delivering high-fidelity human body…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Xinqi Liu , Chenming Wu , Jialun Liu , Xing Liu , Jinbo Wu , Chen Zhao , Haocheng Feng , Errui Ding , Jingdong Wang

Recent advances in neural radiance fields enable novel view synthesis of photo-realistic images in dynamic settings, which can be applied to scenarios with human animation. Commonly used implicit backbones to establish accurate models,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 HyunJun Jung , Nikolas Brasch , Jifei Song , Eduardo Perez-Pellitero , Yiren Zhou , Zhihao Li , Nassir Navab , Benjamin Busam

Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text instructions,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ruikui Wang , Jinheng Feng , Lang Tian , Huaishao Luo , Chaochao Li , Liangbo Zhou , Huan Zhang , Youzheng Wu , Xiaodong He