English
Related papers

Related papers: ZeroAvatar: Zero-shot 3D Avatar Generation from a …

200 papers

Inspired by the effectiveness of 3D Gaussian Splatting (3DGS) in reconstructing detailed 3D scenes within multi-view setups and the emergence of large 2D human foundation models, we introduce Arc2Avatar, the first SDS-based method utilizing…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Dimitrios Gerogiannis , Foivos Paraperas Papantoniou , Rolandos Alexandros Potamias , Alexandros Lattas , Stefanos Zafeiriou

State-of-the-arts text-to-image generation models such as Imagen and Stable Diffusion Model have succeed remarkable progresses in synthesizing high-quality, feature-rich images with high resolution guided by human text prompts. Since…

Computer Vision and Pattern Recognition · Computer Science 2025-01-31 Ziyi Dong , Pengxu Wei , Liang Lin

Self-occlusion is common when capturing people in the wild, where the performer do not follow predefined motion scripts. This challenges existing monocular human reconstruction systems that assume full body visibility. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Zhuoyang Pan , Angjoo Kanazawa , Hang Gao

Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part…

Computer Vision and Pattern Recognition · Computer Science 2021-03-02 Aditya Ramesh , Mikhail Pavlov , Gabriel Goh , Scott Gray , Chelsea Voss , Alec Radford , Mark Chen , Ilya Sutskever

Avatar reconstruction has traditionally relied on per-subject optimization that requires hours of computation or on expensive preprocessing that limits scalability. We introduce FFAvatar, a generalizable feed-forward framework that…

Graphics · Computer Science 2026-05-18 Thuan Hoang Nguyen , Jiahao Luo , Yinyu Nie , Hao Li , Gordon Guocheng Qian , Jian Wang

3D object generation from a single image involves estimating the full 3D geometry and texture of unseen views from an unposed RGB image captured in the wild. Accurately reconstructing an object's complete 3D structure and texture has…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Hritam Basak , Hadi Tabatabaee , Shreekant Gayaka , Ming-Feng Li , Xin Yang , Cheng-Hao Kuo , Arnie Sen , Min Sun , Zhaozheng Yin

Although powerful for image generation, consistent and controllable video is a longstanding problem for diffusion models. Video models require extensive training and computational resources, leading to high costs and large environmental…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Muhammad Haaris Khan , Hadrien Reynaud , Bernhard Kainz

We propose a zero-shot approach to image harmonization, aiming to overcome the reliance on large amounts of synthetic composite images in existing methods. These methods, while showing promising results, involve significant training…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Jianqi Chen , Yilan Zhang , Zhengxia Zou , Keyan Chen , Zhenwei Shi

The problem of modeling an animatable 3D human head avatar under light-weight setups is of significant importance but has not been well solved. Existing 3D representations either perform well in the realism of portrait images synthesis or…

Computer Vision and Pattern Recognition · Computer Science 2023-10-02 Xiaochen Zhao , Lizhen Wang , Jingxiang Sun , Hongwen Zhang , Jinli Suo , Yebin Liu

We propose a neural rendering-based system that creates head avatars from a single photograph. Our approach models a person's appearance by decomposing it into two layers. The first layer is a pose-dependent coarse image that is synthesized…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Egor Zakharov , Aleksei Ivakhnenko , Aliaksandra Shysheya , Victor Lempitsky

Reconstructing 3D humans from a single image has been extensively investigated. However, existing approaches often fall short on capturing fine geometry and appearance details, hallucinating occluded parts with plausible details, and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Zhenzhen Weng , Jingyuan Liu , Hao Tan , Zhan Xu , Yang Zhou , Serena Yeung-Levy , Jimei Yang

Object-level manipulation, relocating or reorienting objects in images or videos while preserving scene realism, is central to film post-production, AR, and creative editing. Yet existing methods struggle to jointly achieve three core…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Penghui Ruan , Bojia Zi , Xianbiao Qi , Youze Huang , Rong Xiao , Pichao Wang , Jiannong Cao , Yuhui Shi

Text-guided diffusion models have shown superior performance in image/video generation and editing. While few explorations have been performed in 3D scenarios. In this paper, we discuss three fundamental and interesting problems on this…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Gang Li , Heliang Zheng , Chaoyue Wang , Chang Li , Changwen Zheng , Dacheng Tao

3D facial avatar reconstruction has been a significant research topic in computer graphics and computer vision, where photo-realistic rendering and flexible controls over poses and expressions are necessary for many related applications.…

Computer Vision and Pattern Recognition · Computer Science 2023-07-10 Wangbo Yu , Yanbo Fan , Yong Zhang , Xuan Wang , Fei Yin , Yunpeng Bai , Yan-Pei Cao , Ying Shan , Yang Wu , Zhongqian Sun , Baoyuan Wu

Recent breakthroughs in text-to-image synthesis have been driven by diffusion models trained on billions of image-text pairs. Adapting this approach to 3D synthesis would require large-scale datasets of labeled 3D data and efficient…

Computer Vision and Pattern Recognition · Computer Science 2022-09-30 Ben Poole , Ajay Jain , Jonathan T. Barron , Ben Mildenhall

Recent advancements in 3D avatar generation excel with multi-view supervision for photorealistic models. However, monocular counterparts lag in quality despite broader applicability. We propose ReCaLaB to close this gap. ReCaLaB is a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Yuchen Rao , Eduardo Perez Pellitero , Benjamin Busam , Yiren Zhou , Jifei Song

Large-scale text-to-video (T2V) diffusion models have great progress in recent years in terms of visual quality, motion and temporal consistency. However, the generation process is still a black box, where all attributes (e.g., appearance,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Jiwen Yu , Xiaodong Cun , Chenyang Qi , Yong Zhang , Xintao Wang , Ying Shan , Jian Zhang

In this paper, we propose a novel approach to reconstruct 3D human body shapes based on a sparse set of RGBD frames using a single RGBD camera. We specifically focus on the realistic settings where human subjects move freely during the…

Computer Vision and Pattern Recognition · Computer Science 2020-06-16 Xinxin Zuo , Sen Wang , Jiangbin Zheng , Weiwei Yu , Minglun Gong , Ruigang Yang , Li Cheng

Creating high-quality 3D avatars using 3D Gaussian Splatting (3DGS) from a monocular video benefits virtual reality and telecommunication applications. However, existing automatic methods exhibit artifacts under novel poses due to limited…

Human-Computer Interaction · Computer Science 2024-12-23 Jotaro Sakamiya , I-Chao Shen , Jinsong Zhang , Mustafa Doga Dogan , Takeo Igarashi

Creating realistic 3D objects and clothed avatars from a single RGB image is an attractive yet challenging problem. Due to its ill-posed nature, recent works leverage powerful prior from 2D diffusion models pretrained on large datasets.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Yuxuan Xue , Xianghui Xie , Riccardo Marin , Gerard Pons-Moll