English
Related papers

Related papers: Tri$^{2}$-plane: Thinking Head Avatar via Feature …

200 papers

Monocular depth estimation aims to infer a dense depth map from a single image, which is a fundamental and prevalent task in computer vision. Many previous works have shown impressive depth estimation results through carefully designed…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Li Liu , Ruijie Zhu , Jiacheng Deng , Ziyang Song , Wenfei Yang , Tianzhu Zhang

We propose a method for synthesizing photo-realistic digital avatars from only one portrait as the reference. Given a portrait, our method synthesizes a coarse talking head video using driving keypoints features. And with the coarse video,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Shaoxu Li

Reconstructing a complete 3D head from a single portrait remains challenging because existing methods still face a sharp quality-speed trade-off: high-fidelity pipelines often rely on multi-stage processing and per-subject optimization,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Yujie Gao , Yao Xiao , Xiangnan Zhu , Ya Li , Yiyi Zhang , Liqing Zhang , Jianfu Zhang

Feedforward monocular face capture methods seek to reconstruct posed faces from a single image of a person. Current state of the art approaches have the ability to regress parametric 3D face models in real-time across a wide range of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Kelian Baert , Shrisha Bharadwaj , Fabien Castan , Benoit Maujean , Marc Christie , Victoria Abrevaya , Adnane Boukhayma

With NeRF widely used for facial reenactment, recent methods can recover photo-realistic 3D head avatar from just a monocular video. Unfortunately, the training process of the NeRF-based methods is quite time-consuming, as MLP used in the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-04 Yuelang Xu , Lizhen Wang , Xiaochen Zhao , Hongwen Zhang , Yebin Liu

This work proposes a novel method to generate realistic talking head videos using audio and visual streams. We animate a source image by transferring head motion from a driving video using a dense motion field generated using learnable…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Madhav Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

High-quality, animatable 3D human avatar reconstruction from monocular videos offers significant potential for reducing reliance on complex hardware, making it highly practical for applications in game development, augmented reality, and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Xia Yuan , Hai Yuan , Wenyi Ge , Ying Fu , Xi Wu , Guanyu Xing

Solving image-to-3D from a single view is an ill-posed problem, and current neural reconstruction methods addressing it through diffusion models still rely on scene-specific optimization, constraining their generalization capability. To…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Christian Simon , Sen He , Juan-Manuel Perez-Rua , Mengmeng Xu , Amine Benhalloum , Tao Xiang

We present a 3D-aware one-shot head reenactment method based on a fully volumetric neural disentanglement framework for source appearance and driver expressions. Our method is real-time and produces high-fidelity and view-consistent output,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Phong Tran , Egor Zakharov , Long-Nhat Ho , Anh Tuan Tran , Liwen Hu , Hao Li

We introduce VOODOO XP: a 3D-aware one-shot head reenactment method that can generate highly expressive facial expressions from any input driver video and a single 2D portrait. Our solution is real-time, view-consistent, and can be…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Phong Tran , Egor Zakharov , Long-Nhat Ho , Liwen Hu , Adilbek Karmanov , Aviral Agarwal , McLean Goldwhite , Ariana Bermudez Venegas , Anh Tuan Tran , Hao Li

Convolutional neural network (CNN) slides a kernel over the whole image to produce an output map. This kernel scheme reduces the number of parameters with respect to a fully connected neural network (NN). While CNN has proven to be an…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Ihsan Ullah , Alfredo Petrosino

Efficiently reconstructing 3D scenes from monocular video remains a core challenge in computer vision, vital for applications in virtual reality, robotics, and scene understanding. Recently, frame-by-frame progressive reconstruction without…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Wenyan Cong , Hanqing Zhu , Kevin Wang , Jiahui Lei , Colton Stearns , Yuanhao Cai , Leonidas Guibas , Zhangyang Wang , Zhiwen Fan

In recent years, 3D models have gained popularity in various fields, including entertainment, manufacturing, and simulation. However, manually creating these models can be a time-consuming and resource-intensive process, making it…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Potito Aghilar , Vito Walter Anelli , Michelantonio Trizio , Tommaso Di Noia

Reconstructing articulated 3D objects from a single image requires jointly inferring object geometry, part structure, and motion parameters from limited visual evidence. A key difficulty lies in the entanglement between motion cues and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Haitian Li , Haozhe Xie , Junxiang Xu , Beichen Wen , Fangzhou Hong , Ziwei Liu

Attention-based models such as transformers have shown outstanding performance on dense prediction tasks, such as semantic segmentation, owing to their capability of capturing long-range dependency in an image. However, the benefit of…

Computer Vision and Pattern Recognition · Computer Science 2022-07-13 Ashutosh Agarwal , Chetan Arora

Despite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Wei Yin , Jianming Zhang , Oliver Wang , Simon Niklaus , Simon Chen , Yifan Liu , Chunhua Shen

We propose the Multiple View Performer (MVP) - a new architecture for 3D shape completion from a series of temporally sequential views. MVP accomplishes this task by using linear-attention Transformers called Performers. Our model allows…

Computer Vision and Pattern Recognition · Computer Science 2023-02-13 David Watkins , Peter Allen , Krzysztof Choromanski , Jacob Varley , Nicholas Waytowich

3D face reconstruction and face alignment are two fundamental and highly related topics in computer vision. Recently, some works start to use deep learning models to estimate the 3DMM coefficients to reconstruct 3D face geometry. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Zihao Jian , Minshan Xie

Neural Parametric Head Models (NPHMs) are a recent advancement over mesh-based 3d morphable models (3DMMs) to facilitate high-fidelity geometric detail. However, fitting NPHMs to visual inputs is notoriously challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Simon Giebenhain , Tobias Kirschstein , Liam Schoneveld , Davide Davoli , Zhe Chen , Matthias Nießner

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Shuling Zhao , Fa-Ting Hong , Xiaoshui Huang , Dan Xu