English
Related papers

Related papers: Learning to Fuse 2D and 3D Image Cues for Monocula…

200 papers

Most 3d human pose estimation methods assume that input -- be it images of a scene collected from one or several viewpoints, or from a video -- is given. Consequently, they focus on estimates leveraging prior knowledge and measurement by…

Computer Vision and Pattern Recognition · Computer Science 2020-12-17 Erik Gärtner , Aleksis Pirinen , Cristian Sminchisescu

Despite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Wei Yin , Jianming Zhang , Oliver Wang , Simon Niklaus , Simon Chen , Yifan Liu , Chunhua Shen

Depictions of similar human body configurations can vary with changing viewpoints. Using only 2D information, we would like to enable vision algorithms to recognize similarity in human body poses across multiple views. This ability is…

Computer Vision and Pattern Recognition · Computer Science 2020-10-26 Jennifer J. Sun , Jiaping Zhao , Liang-Chieh Chen , Florian Schroff , Hartwig Adam , Ting Liu

In 3D human pose estimation one of the biggest problems is the lack of large, diverse datasets. This is especially true for multi-person 3D pose estimation, where, to our knowledge, there are only machine generated annotations available for…

Computer Vision and Pattern Recognition · Computer Science 2020-04-09 Marton Veges , Andras Lorincz

Current unsupervised 2D-3D human pose estimation (HPE) methods do not work in multi-person scenarios due to perspective ambiguity in monocular images. Therefore, we present one of the first studies investigating the feasibility of…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Peter Hardy , Hansung Kim

Estimating the 3D poses of hands and objects from a single RGB image is a fundamental yet challenging problem, with broad applications in augmented reality and human-computer interaction. Existing methods largely rely on visual cues alone,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Jun Zhou , Chi Xu , Kaifeng Tang , Yuting Ge , Tingrui Guo , Li Cheng

Either RGB images or inertial signals have been used for the task of motion capture (mocap), but combining them together is a new and interesting topic. We believe that the combination is complementary and able to solve the inherent…

Computer Vision and Pattern Recognition · Computer Science 2023-09-04 Shaohua Pan , Qi Ma , Xinyu Yi , Weifeng Hu , Xiong Wang , Xingkang Zhou , Jijunnan Li , Feng Xu

Monocular 3D scene understanding tasks, such as object size estimation, heading angle estimation and 3D localization, is challenging. Successful modern day methods for 3D scene understanding require the use of a 3D sensor. On the other…

Computer Vision and Pattern Recognition · Computer Science 2019-09-04 Xinshuo Weng , Kris Kitani

Monocular 3D human pose and shape estimation is an inherently ill-posed problem due to depth ambiguities, occlusions, and truncations. Recent probabilistic approaches learn a distribution over plausible 3D human meshes by maximizing the…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Tom Wehrbein , Marco Rudolph , Bodo Rosenhahn , Bastian Wandt

Human bodies exhibit various shapes for different identities or poses, but the body shape has certain similarities in structure and thus can be embedded in a low-dimensional space. This paper presents an autoencoder-like network…

Computer Vision and Pattern Recognition · Computer Science 2020-04-20 Boyi Jiang , Juyong Zhang , Jianfei Cai , Jianmin Zheng

This paper presents a novel 3D human pose estimation approach using a single stream of asynchronous events as input. Most of the state-of-the-art approaches solve this task with RGB cameras, however struggling when subjects are moving fast.…

Computer Vision and Pattern Recognition · Computer Science 2021-04-22 Gianluca Scarpellini , Pietro Morerio , Alessio Del Bue

Recent advances in data-driven geometric multi-view 3D reconstruction foundation models (e.g., DUSt3R) have shown remarkable performance across various 3D vision tasks, facilitated by the release of large-scale, high-quality 3D datasets.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Wenyu Li , Sidun Liu , Peng Qiao , Yong Dou

In this paper, we propose a pose grammar to tackle the problem of 3D human pose estimation. Our model directly takes 2D pose as input and learns a generalized 2D-3D mapping function. The proposed model consists of a base network which…

Computer Vision and Pattern Recognition · Computer Science 2018-01-08 Haoshu Fang , Yuanlu Xu , Wenguan Wang , Xiaobai Liu , Song-Chun Zhu

Recently, the vision transformer and its variants have played an increasingly important role in both monocular and multi-view human pose estimation. Considering image patches as tokens, transformers can model the global dependencies within…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Haoyu Ma , Zhe Wang , Yifei Chen , Deying Kong , Liangjian Chen , Xingwei Liu , Xiangyi Yan , Hao Tang , Xiaohui Xie

Learning methods for relative camera pose estimation have been developed largely in isolation from classical geometric approaches. The question of how to integrate predictions from deep neural networks (DNNs) and solutions from geometric…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Bingbing Zhuang , Manmohan Chandraker

We introduce a new approach for estimating the 3D pose and the 3D shape of an object from a single image. Given a training set of view exemplars, we learn and select appearance-based discriminative parts which are mapped onto the 3D model…

Computer Vision and Pattern Recognition · Computer Science 2015-02-03 Menglong Zhu , Xiaowei Zhou , Kostas Daniilidis

Recognition of human poses and actions is crucial for autonomous systems to interact smoothly with people. However, cameras generally capture human poses in 2D as images and videos, which can have significant appearance variations across…

Computer Vision and Pattern Recognition · Computer Science 2021-11-19 Ting Liu , Jennifer J. Sun , Long Zhao , Jiaping Zhao , Liangzhe Yuan , Yuxiao Wang , Liang-Chieh Chen , Florian Schroff , Hartwig Adam

The common approach to 3D human pose estimation is predicting the body joint coordinates relative to the hip. This works well for a single person but is insufficient in the case of multiple interacting people. Methods predicting absolute…

Computer Vision and Pattern Recognition · Computer Science 2019-04-15 Márton Véges , András Lőrincz

For human pose estimation in monocular images, joint occlusions and overlapping upon human bodies often result in deviated pose predictions. Under these circumstances, biologically implausible pose predictions may be produced. In contrast,…

Computer Vision and Pattern Recognition · Computer Science 2017-05-03 Yu Chen , Chunhua Shen , Xiu-Shen Wei , Lingqiao Liu , Jian Yang

A monocular 3D object tracking system generally has only up-to-scale pose estimation results without any prior knowledge of the tracked object. In this paper, we propose a novel idea to recover the metric scale of an arbitrary dynamic…

Robotics · Computer Science 2018-08-22 Kejie Qiu , Tong Qin , Hongwen Xie , Shaojie Shen