English
Related papers

Related papers: HART: Human Aligned Reconstruction Transformer

200 papers

We consider the problem of estimating a parametric model of 3D human mesh from a single image. While there has been substantial recent progress in this area with direct regression of model parameters, these methods only implicitly exploit…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Georgios Georgakis , Ren Li , Srikrishna Karanam , Terrence Chen , Jana Kosecka , Ziyan Wu

Previous works on Human Pose and Shape Estimation (HPSE) from RGB images can be broadly categorized into two main groups: parametric and non-parametric approaches. Parametric techniques leverage a low-dimensional statistical body model for…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Guénolé Fiche , Simon Leglaive , Xavier Alameda-Pineda , Antonio Agudo , Francesc Moreno-Noguer

We present Multi-HMR, a strong sigle-shot model for multi-person 3D human mesh recovery from a single RGB image. Predictions encompass the whole body, i.e., including hands and facial expressions, using the SMPL-X parametric model and 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Fabien Baradel , Matthieu Armando , Salma Galaaoui , Romain Brégier , Philippe Weinzaepfel , Grégory Rogez , Thomas Lucas

No augmented application is possible without animated humanoid avatars. At the same time, generating human replicas from real-world monocular hand-held or robotic sensor setups is challenging due to the limited availability of views.…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Alessandro Sanvito , Andrea Ramazzina , Stefanie Walz , Mario Bijelic , Felix Heide

Modeling animatable human avatars from videos is a long-standing and challenging problem. While conventional methods require per-instance optimization, recent feed-forward methods have been proposed to generate 3D Gaussians with a learnable…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Yifan Liu , Shengjun Zhang , Chensheng Dai , Yang Chen , Hao Liu , Chen Li , Yueqi Duan

Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure…

Computer Vision and Pattern Recognition · Computer Science 2023-11-01 Junkun Yuan , Xinyu Zhang , Hao Zhou , Jian Wang , Zhongwei Qiu , Zhiyin Shao , Shaofeng Zhang , Sifan Long , Kun Kuang , Kun Yao , Junyu Han , Errui Ding , Lanfen Lin , Fei Wu , Jingdong Wang

Neural reconstruction and rendering strategies have demonstrated state-of-the-art performances due, in part, to their ability to preserve high level shape details. Existing approaches, however, either represent objects as implicit surface…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Angtian Wang , Yuanlu Xu , Nikolaos Sarafianos , Robert Maier , Edmond Boyer , Alan Yuille , Tony Tung

The field of 3D detailed human mesh reconstruction has made significant progress in recent years. However, current methods still face challenges when used in industrial applications due to unstable results, low-quality meshes, and a lack of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Xiaoyu Zhan , Jianxin Yang , Yuanqi Li , Jie Guo , Yanwen Guo , Wenping Wang

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Lingteng Qiu , Xiaodong Gu , Peihao Li , Qi Zuo , Weichao Shen , Junfei Zhang , Kejie Qiu , Weihao Yuan , Guanying Chen , Zilong Dong , Liefeng Bo

We introduce Gaussian Articulated Template Model GART, an explicit, efficient, and expressive representation for non-rigid articulated subject capturing and rendering from monocular videos. GART utilizes a mixture of moving 3D Gaussians to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Jiahui Lei , Yufu Wang , Georgios Pavlakos , Lingjie Liu , Kostas Daniilidis

Animating realistic character interactions with the surrounding environment is important for autonomous agents in gaming, AR/VR, and robotics. However, current methods for human motion reconstruction struggle with accurately placing humans…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Joshua Li , Brendan Chharawala , Chang Shu , Xue Bin Peng , Pengcheng Xi

We present Splat-SAP, a feed-forward approach to render novel views of human-centered scenes from binocular cameras with large sparsity. Gaussian Splatting has shown its promising potential in rendering tasks, but it typically necessitates…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Boyao Zhou , Shunyuan Zheng , Zhanfeng Liao , Zihan Ma , Hanzhang Tu , Boning Liu , Yebin Liu

Most models of visual attention aim at predicting either top-down or bottom-up control, as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT), a single model that…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Zhibo Yang , Sounak Mondal , Seoyoung Ahn , Ruoyu Xue , Gregory Zelinsky , Minh Hoai , Dimitris Samaras

This work addresses the problem of real-time rendering of photorealistic human body avatars learned from multi-view videos. While the classical approaches to model and render virtual humans generally use a textured mesh, recent research has…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Arthur Moreau , Jifei Song , Helisa Dhamo , Richard Shaw , Yiren Zhou , Eduardo Pérez-Pellitero

In this paper, we introduce a method for reconstructing 3D humans from a single image using a biomechanically accurate skeleton model. To achieve this, we train a transformer that takes an image as input and estimates the parameters of the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Yan Xia , Xiaowei Zhou , Etienne Vouga , Qixing Huang , Georgios Pavlakos

We propose an approach for optimizing high-quality clothed human body shapes in minutes, using multi-view posed images. While traditional neural rendering methods struggle to disentangle geometry and appearance using only rendering loss,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Lixiang Lin , Songyou Peng , Qijun Gan , Jianke Zhu

Reconstructing posed 3D human models from monocular images has important applications in the sports industry, including performance tracking, injury prevention and virtual training. In this work, we combine 3D human pose and shape…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Lorenza Prospero , Abdullah Hamdi , Joao F. Henriques , Christian Rupprecht

Surface reconstruction has been widely studied in computer vision and graphics. However, existing surface reconstruction works struggle to recover accurate scene geometry when the input views are extremely sparse. To address this issue, we…

Graphics · Computer Science 2025-11-26 Hanzhi Chang , Ruijie Zhu , Wenjie Chang , Mulin Yu , Yanzhe Liang , Jiahao Lu , Zhuoyuan Li , Tianzhu Zhang

We introduce Hybrid Autoregressive Transformer (HART), an autoregressive (AR) visual generation model capable of directly generating 1024x1024 images, rivaling diffusion models in image generation quality. Existing AR models face…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Haotian Tang , Yecheng Wu , Shang Yang , Enze Xie , Junsong Chen , Junyu Chen , Zhuoyang Zhang , Han Cai , Yao Lu , Song Han

Recovering high-fidelity 3D hand geometry from images is a critical task in computer vision, holding significant value for domains such as robotics, animation and VR/AR. Crucially, scalable applications demand both accuracy and deployment…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yumeng Liu , Xiao-Xiao Long , Marc Habermann , Xuanze Yang , Cheng Lin , Yuan Liu , Yuexin Ma , Wenping Wang , Ligang Liu