English
Related papers

Related papers: SMPLer: Taming Transformers for Monocular 3D Human…

200 papers

This paper addresses the problem of 3D human pose and shape estimation from a single image. Previous approaches consider a parametric model of the human body, SMPL, and attempt to regress the model parameters that give rise to a mesh…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Nikos Kolotouros , Georgios Pavlakos , Kostas Daniilidis

Human Mesh Recovery (HMR) is an important yet challenging problem with applications across various domains including motion capture, augmented reality, and biomechanics. Accurately predicting human pose parameters from a single image…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Jaewoo Heo , George Hu , Zeyu Wang , Serena Yeung-Levy

Existing multi-person pose estimators can be roughly divided into two-stage approaches (top-down and bottom-up approaches) and one-stage approaches. The two-stage methods either suffer high computational redundancy for additional person…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Junqi Lin , Huixin Miao , Junjie Cao , Zhixun Su , Risheng Liu

Reconstructing detailed 3D human meshes from a single in-the-wild image remains a fundamental challenge in computer vision. Existing SMPLX-based methods often suffer from slow inference, produce only coarse body poses, and exhibit…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Jiahao Wu , Yunfei Liu , Lijian Lin , Ye Zhu , Lei Zhu , Jingyi Li , Yu Li

3D human pose estimation can be handled by encoding the geometric dependencies between the body parts and enforcing the kinematic constraints. Recently, Transformer has been adopted to encode the long-range dependencies between the joints…

Computer Vision and Pattern Recognition · Computer Science 2022-03-28 Mohammed Hassanin , Abdelwahed Khamiss , Mohammed Bennamoun , Farid Boussaid , Ibrahim Radwan

The 2D human pose estimation (HPE) is a basic visual problem. However, its supervised learning requires massive keypoint labels, which is labor-intensive to collect. Thus, we aim at boosting a pose estimator by excavating extra unlabeled…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Huayi Zhou , Mukun Luo , Fei Jiang , Yue Ding , Hongtao Lu , Kui Jia

Great progress has been made in estimating 3D human pose and shape from images and video by training neural networks to directly regress the parameters of parametric human models like SMPL. However, existing body models have simplified…

Graphics · Computer Science 2025-09-09 Marilyn Keller , Keenon Werling , Soyong Shin , Scott Delp , Sergi Pujades , C. Karen Liu , Michael J. Black

Photorealistic 3D full-body human reconstruction from a single image is a critical yet challenging task for applications in films and video games due to inherent ambiguities and severe self-occlusions. While recent approaches leverage SMPL…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Wenyue Chen , Peng Li , Wangguandong Zheng , Chengfeng Zhao , Mengfei Li , Yaolong Zhu , Zhiyang Dou , Ronggang Wang , Yuan Liu

Transformer models rely on self-attention to capture token dependencies but face challenges in effectively integrating positional information while allowing multi-head attention (MHA) flexibility. Prior methods often model semantic and…

Machine Learning · Computer Science 2025-05-28 Jintian Shao , Hongyi Huang , Jiayi Wu , Beiwen Zhang , ZhiYu Wu , You Shan , MingKai Zheng

Transformers have shown great potential in computer vision tasks. A common belief is their attention-based token mixer module contributes most to their competence. However, recent works show the attention-based module in Transformers can be…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Weihao Yu , Mi Luo , Pan Zhou , Chenyang Si , Yichen Zhou , Xinchao Wang , Jiashi Feng , Shuicheng Yan

We propose the Multiple View Performer (MVP) - a new architecture for 3D shape completion from a series of temporally sequential views. MVP accomplishes this task by using linear-attention Transformers called Performers. Our model allows…

Computer Vision and Pattern Recognition · Computer Science 2023-02-13 David Watkins , Peter Allen , Krzysztof Choromanski , Jacob Varley , Nicholas Waytowich

We describe an end-to-end method for recovering 3D human body mesh from single images and monocular videos. Different from the existing methods try to obtain all the complex 3D pose, shape, and camera parameters from one coupling feature,…

Computer Vision and Pattern Recognition · Computer Science 2019-09-18 Sun Yu , Ye Yun , Liu Wu , Gao Wenpeng , Fu YiLi , Mei Tao

Transformer encoder architectures have recently achieved state-of-the-art results on monocular 3D human mesh reconstruction, but they require a substantial number of parameters and expensive computations. Due to the large memory overhead…

Computer Vision and Pattern Recognition · Computer Science 2022-07-29 Junhyeong Cho , Kim Youwang , Tae-Hyun Oh

We address the challenges in estimating 3D human poses from multiple views under occlusion and with limited overlapping views. We approach multi-view, single-person 3D human pose reconstruction as a regression problem and propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Olivier Moliner , Sangxia Huang , Kalle Åström

Generating expressive conducting gestures from music is a challenging cross-modal motion synthesis problem: the output must follow long-range musical structure, preserve beat-level synchronization, and remain plausible as a fine-grained 3D…

Sound · Computer Science 2026-05-05 Ke Qiu , Yawen Qin , Tianzhi Jia , Xiaole Yang , Kaimin Wang , Kaixing Yang

A common approach to localize 3D human joints in a synchronized and calibrated multi-view setup consists of two-steps: (1) apply a 2D detector separately on each view to localize joints in 2D, and (2) perform robust triangulation on 2D…

Computer Vision and Pattern Recognition · Computer Science 2020-05-12 Yihui He , Rui Yan , Katerina Fragkiadaki , Shoou-I Yu

Current state-of-the-art methods cast monocular 3D human pose estimation as a learning problem by training neural networks on large data sets of images and corresponding skeleton poses. In contrast, we propose an approach that can exploit…

Computer Vision and Pattern Recognition · Computer Science 2020-10-14 Simon Jenni , Paolo Favaro

We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle to recover metric scale due to monocular scale ambiguity and weak-perspective camera…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Chentao Song , He Zhang , Haolei Yuan , Haozhe Lin , Jianhua Tao , Hongwen Zhang , Tao Yu

We propose a novel method for joint estimation of shape and pose of rigid objects from their sequentially observed RGB-D images. In sharp contrast to past approaches that rely on complex non-linear optimization, we propose to formulate it…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Yuta Yoshitake , Mai Nishimura , Shohei Nobuhara , Ko Nishino

We present SLOPER4D, a novel scene-aware dataset collected in large urban environments to facilitate the research of global human pose estimation (GHPE) with human-scene interaction in the wild. Employing a head-mounted device integrated…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Yudi Dai , Yitai Lin , Xiping Lin , Chenglu Wen , Lan Xu , Hongwei Yi , Siqi Shen , Yuexin Ma , Cheng Wang
‹ Prev 1 3 4 5 6 7 10 Next ›