English
Related papers

Related papers: Back to Optimization: Diffusion-based Zero-Shot 3D…

200 papers

Fully-supervised category-level pose estimation aims to determine the 6-DoF poses of unseen instances from known categories, requiring expensive mannual labeling costs. Recently, various self-supervised category-level pose estimation…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Jingtao Sun , Yaonan Wang , Mingtao Feng , Chao Ding , Mike Zheng Shou , Ajmal Saeed Mian

3D human pose estimation (3D HPE) has emerged as a prominent research topic, particularly in the realm of RGB-based methods. However, the use of RGB images is often limited by issues such as occlusion and privacy constraints. Consequently,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Mengshi Qi , Jiaxuan Peng , Xianlin Zhang , Huadong Ma

We tackle the problem of Human Mesh Recovery (HMR) from a single RGB image, formulating it as an image-conditioned human pose and shape generation. While recovering 3D human pose from 2D observations is inherently ambiguous, most existing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Donghwan Kim , Tae-Kyun Kim

Current unsupervised 2D-3D human pose estimation (HPE) methods do not work in multi-person scenarios due to perspective ambiguity in monocular images. Therefore, we present one of the first studies investigating the feasibility of…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Peter Hardy , Hansung Kim

Due to the difficulty of acquiring large-scale 3D human keypoint annotation, previous methods for 3D human pose estimation (HPE) have often relied on 2D image features and sequential 2D annotations. Furthermore, the training of these…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Dongqiangzi Ye , Yufei Xie , Weijia Chen , Zixiang Zhou , Lingting Ge , Hassan Foroosh

Monocular 3D human pose estimation remains a challenging task due to inherent depth ambiguities and occlusions. Compared to traditional methods based on Transformers or Convolutional Neural Networks (CNNs), recent diffusion-based approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Haoxin Yang , Weihong Chen , Xuemiao Xu , Cheng Xu , Peng Xiao , Cuifeng Sun , Shaoyu Huang , Shengfeng He

This work targets to construct a robust human pose prior. However, it remains a persistent challenge due to biomechanical constraints and diverse human movements. Traditional priors like VAEs and NDFs often exhibit shortcomings in realism…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Junzhe Lu , Jing Lin , Hongkun Dou , Ailing Zeng , Yue Deng , Yulun Zhang , Haoqian Wang

Following the success of deep convolutional networks, state-of-the-art methods for 3d human pose estimation have focused on deep end-to-end systems that predict 3d joint locations given raw image pixels. Despite their excellent performance,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-08 Julieta Martinez , Rayat Hossain , Javier Romero , James J. Little

Recovering 3D human poses from a monocular camera view is a highly ill-posed problem due to the depth ambiguity. Earlier studies on 3D human pose lifting from 2D often contain incorrect-yet-overconfident 3D estimations. To mitigate the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Cuong Le , Pavlo Melnyk , Bastian Wandt , Mårten Wadenbäck

Accurate and real-time three-dimensional (3D) pose estimation is challenging in resource-constrained and dynamic environments owing to its high computational complexity. To address this issue, this study proposes a novel cooperative…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Hyun-Ho Choi , Kangsoo Kim , Ki-Ho Lee , Kisong Lee

Denoising diffusion probabilistic models that were initially proposed for realistic image generation have recently shown success in various perception tasks (e.g., object detection and image segmentation) and are increasingly gaining…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Runyang Feng , Yixing Gao , Tze Ho Elden Tse , Xueqing Ma , Hyung Jin Chang

3D human pose and shape estimation (a.k.a. "human mesh recovery") has achieved substantial progress. Researchers mainly focus on the development of novel algorithms, while less attention has been paid to other critical factors involved.…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Hui En Pang , Zhongang Cai , Lei Yang , Tianwei Zhang , Ziwei Liu

Pre-training is a general method that is used in a range of deep learning tasks. By first training a model on one task, and then further training on the downstream task used for final evaluation, the model is forced to learn a more general…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Liyao Jiang , Ruichen Chen , Keith G. Mills

In-the-wild human pose estimation has a huge potential for various fields, ranging from animation and action recognition to intention recognition and prediction for autonomous driving. The current state-of-the-art is focused only on RGB and…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Michael Fürst , Shriya T. P. Gupta , René Schuster , Oliver Wasenmüller , Didier Stricker

Training accurate 3D human pose estimators requires large amount of 3D ground-truth data which is costly to collect. Various weakly or self supervised pose estimation methods have been proposed due to lack of 3D data. Nevertheless, these…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Muhammed Kocabas , Salih Karagoz , Emre Akbas

3D Human body pose and shape estimation within a temporal sequence can be quite critical for understanding human behavior. Despite the significant progress in human pose estimation in the recent years, which are often based on single images…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Zhouping Wang , Sarah Ostadabbas

In recent times, there has been a growing interest in developing effective perception techniques for combining information from multiple modalities. This involves aligning features obtained from diverse sources to enable more efficient…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zhongyu Jiang , Wenhao Chai , Lei Li , Zhuoran Zhou , Cheng-Yen Yang , Jenq-Neng Hwang

Traditional methods for human localization and pose estimation (HPE), which mainly rely on RGB images as an input modality, confront substantial limitations in real-world applications due to privacy concerns. In contrast, radar-based HPE…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Yuan-Hao Ho , Jen-Hao Cheng , Sheng Yao Kuan , Zhongyu Jiang , Wenhao Chai , Hsiang-Wei Huang , Chih-Lung Lin , Jenq-Neng Hwang

Monocular 3D pose estimation is fundamentally ill-posed due to depth ambiguity and occlusions, thereby motivating probabilistic methods that generate multiple plausible 3D pose hypotheses. In particular, diffusion-based models have recently…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Ti Wang , Xiaohang Yu , Mackenzie Weygandt Mathis

Direct Preference Optimization (DPO) has been successfully used to align large language models (LLMs) according to human preferences, and more recently it has also been applied to improving the quality of text-to-image diffusion models.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Shivanshu Shekhar , Shreyas Singh , Tong Zhang