English
Related papers

Related papers: ICo3D: An Interactive Conversational 3D Virtual Hu…

200 papers

Existing full-body Gaussian avatar methods primarily optimize global reconstruction quality and often fail to preserve fine-grained facial geometry and expression details. This challenge arises from limited facial representational capacity…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Willem Menu , Erkut Akdag , Pedro Quesado , Yasaman Kashefbahrami , Egor Bondarev

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Sai Kumar Dwivedi , Dimitrije Antić , Shashank Tripathi , Omid Taheri , Cordelia Schmid , Michael J. Black , Dimitrios Tzionas

We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Qianyun He , Xinya Ji , Yicheng Gong , Yuanxun Lu , Zhengyu Diao , Linjia Huang , Yao Yao , Siyu Zhu , Zhan Ma , Songcen Xu , Xiaofei Wu , Zixiao Zhang , Xun Cao , Hao Zhu

With the booming of virtual reality (VR) technology, there is a growing need for customized 3D avatars. However, traditional methods for 3D avatar modeling are either time-consuming or fail to retain similarity to the person being modeled.…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Chuanyu Pan , Guowei Yang , Taijiang Mu , Yu-Kun Lai

With the rapid advancement of 3D representation techniques and generative models, substantial progress has been made in reconstructing full-body 3D avatars from a single image. However, this task remains fundamentally ill-posedness due to…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Gaofeng Liu , Hengsen Li , Ruoyu Gao , Xuetong Li , Zhiyuan Ma , Tao Fang

We introduce 3D Gaussian blendshapes for modeling photorealistic head avatars. Taking a monocular video as input, we learn a base head model of neutral expression, along with a group of expression blendshapes, each of which corresponds to a…

Graphics · Computer Science 2024-05-03 Shengjie Ma , Yanlin Weng , Tianjia Shao , Kun Zhou

While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's speaking status…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Yasheng Sun , Wenqing Chu , Hang Zhou , Kaisiyuan Wang , Hideki Koike

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xuangeng Chu , Yuan Gan , Ziteng Cui , Shuhong Liu , Jian Wang , Bing Zhou , Tatsuya Harada

Numerous methods have been proposed to detect, estimate, and analyze properties of people in images, including 3D pose, shape, contact, human-object interaction, and emotion. While widely applicable in vision and other areas, such methods…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Jing Lin , Yao Feng , Weiyang Liu , Michael J. Black

Generating animatable human avatars from a single image is essential for various digital human modeling applications. Existing 3D reconstruction methods often struggle to capture fine details in animatable models, while generative…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Lingteng Qiu , Shenhao Zhu , Qi Zuo , Xiaodong Gu , Yuan Dong , Junfei Zhang , Chao Xu , Zhe Li , Weihao Yuan , Liefeng Bo , Guanying Chen , Zilong Dong

Audio-driven 3D talking avatar generation is increasingly important in virtual communication, digital humans, and interactive media, where avatars must preserve identity, synchronize lip motion with speech, express emotion, and exhibit…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Zhongju Wang , Zhenhong Sun , Beier Wang , Yifu Wang , Daoyi Dong , Huadong Mo , Hongdong Li

To address the ill-posed problem caused by partial observations in monocular human volumetric capture, we present AvatarCap, a novel framework that introduces animatable avatars into the capture pipeline for high-fidelity reconstruction in…

Computer Vision and Pattern Recognition · Computer Science 2022-07-13 Zhe Li , Zerong Zheng , Hongwen Zhang , Chaonan Ji , Yebin Liu

This work presents an audio-visual interactive chatbot (AVIN-Chat) system that allows users to have face-to-face conversations with 3D avatars in real-time. Compared to the previous chatbot services, which provide text-only or speech-only…

Human-Computer Interaction · Computer Science 2024-09-04 Chanhyuk Park , Jungbin Cho , Junwan Kim , Seongmin Lee , Jungsu Kim , Sanghoon Lee

Existing methods for image-to-3D avatar generation struggle to produce highly detailed, animation-ready avatars suitable for real-world applications. We introduce AdaHuman, a novel framework that generates high-fidelity animatable 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Yangyi Huang , Ye Yuan , Xueting Li , Jan Kautz , Umar Iqbal

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bin Liu , Zhixiang Xiong , Zhifen He , Bo Li

We introduce a new conversation head generation benchmark for synthesizing behaviors of a single interlocutor in a face-to-face conversation. The capability to automatically synthesize interlocutors which can participate in long and…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Mohan Zhou , Yalong Bai , Wei Zhang , Ting Yao , Tiejun Zhao

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Zhenhui Ye , Tianyun Zhong , Yi Ren , Jiaqi Yang , Weichuang Li , Jiawei Huang , Ziyue Jiang , Jinzheng He , Rongjie Huang , Jinglin Liu , Chen Zhang , Xiang Yin , Zejun Ma , Zhou Zhao

The creation of 3D human avatars from multi-view videos is a significant yet challenging task in computer vision. However, existing techniques rely on high-quality, sharp images as input, which are often impractical to obtain in real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Muyao Niu , Yifan Zhan , Qingtian Zhu , Zhuoxiao Li , Wei Wang , Zhihang Zhong , Xiao Sun , Yinqiang Zheng

Constructing vivid 3D head avatars for given subjects and realizing a series of animations on them is valuable yet challenging. This paper presents GaussianHead, which models the actional human head with anisotropic 3D Gaussians. In our…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Jie Wang , Jiu-Cheng Xie , Xianyan Li , Feng Xu , Chi-Man Pun , Hao Gao

This work addresses the problem of real-time rendering of photorealistic human body avatars learned from multi-view videos. While the classical approaches to model and render virtual humans generally use a textured mesh, recent research has…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Arthur Moreau , Jifei Song , Helisa Dhamo , Richard Shaw , Yiren Zhou , Eduardo Pérez-Pellitero