English
Related papers

Related papers: Lifting Motion to the 3D World via 2D Diffusion

200 papers

Inferring 3D human motion from video remains a challenging problem with many applications. While traditional methods estimate the human in image coordinates, many applications require human motion to be estimated in world coordinates. This…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Joachim Tesch , Giorgio Becherini , Prerana Achar , Anastasios Yiannakidis , Muhammed Kocabas , Priyanka Patel , Michael J. Black

Recent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image-conditioned models. It is inconvenient for them to take advantage of various off-the-shelf 3D assets with multi-view…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Yanqin Jiang , Chaohui Yu , Chenjie Cao , Fan Wang , Weiming Hu , Jin Gao

From an image of a person in action, we can easily guess the 3D motion of the person in the immediate past and future. This is because we have a mental model of 3D human dynamics that we have acquired from observing visual sequences of…

Computer Vision and Pattern Recognition · Computer Science 2019-09-18 Angjoo Kanazawa , Jason Y. Zhang , Panna Felsen , Jitendra Malik

Generating consistent multiple views for 3D reconstruction tasks is still a challenge to existing image-to-3D diffusion models. Generally, incorporating 3D representations into diffusion model decrease the model's speed as well as…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Emmanuelle Bourigault , Pauline Bourigault

After many researchers observed fruitfulness from the recent diffusion probabilistic model, its effectiveness in image generation is actively studied these days. In this paper, our objective is to evaluate the potential of diffusion…

Computer Vision and Pattern Recognition · Computer Science 2023-03-01 Hyemin Ahn , Esteve Valls Mascaro , Dongheui Lee

Recent open-world 3D representation learning methods using Vision-Language Models (VLMs) to align 3D point cloud with image-text information have shown superior 3D zero-shot performance. However, CAD-rendered images for this alignment often…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Ye Mao , Junpeng Jing , Krystian Mikolajczyk

In this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Yatian Pang , Bin Zhu , Bin Lin , Mingzhe Zheng , Francis E. H. Tay , Ser-Nam Lim , Harry Yang , Li Yuan

Estimating 3D poses of multiple humans in real-time is a classic but still challenging task in computer vision. Its major difficulty lies in the ambiguity in cross-view association of 2D poses and the huge state space when there are…

Computer Vision and Pattern Recognition · Computer Science 2021-07-30 Long Chen , Haizhou Ai , Rui Chen , Zijie Zhuang , Shuang Liu

3D human pose lifting from a single RGB image is a challenging task in 3D vision. Existing methods typically establish a direct joint-to-joint mapping from 2D to 3D poses based on 2D features. This formulation suffers from two fundamental…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Jinghong Zheng , Changlong Jiang , Yang Xiao , Jiaqi Li , Haohong Kuang , Hang Xu , Ran Wang , Zhiguo Cao , Min Du , Joey Tianyi Zhou

Monocular 3D human pose estimation (HPE) often encounters challenges such as depth ambiguity and occlusion during the 2D-to-3D lifting process. Additionally, traditional methods may overlook multi-scale skeleton features when utilizing…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Bing Han , Yuhua Huang , Pan Gao

We propose a method of estimating a 3D human pose from a single view without 3D supervision. The key to our method is to leverage the 2D diffusion priors of motion diffusion models (MDMs) pre-trained on large 2D human pose datasets.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Ryohei Goto , Takuya Fujihashi , Shunsuke Saruwatari , Fumio Okura

Extracting keypoint locations from input hand frames, known as 3D hand pose estimation, is a critical task in various human-computer interaction applications. Essentially, the 3D hand pose estimation can be regarded as a 3D point subset…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Wencan Cheng , Hao Tang , Luc Van Gool , Jong Hwan Ko

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has increasingly focused on…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yueru Jia , Jiaming Liu , Sixiang Chen , Chenyang Gu , Zhilue Wang , Longzan Luo , Lily Lee , Pengwei Wang , Zhongyuan Wang , Renrui Zhang , Shanghang Zhang

Predicting 3D human poses in real-world scenarios, also known as human pose forecasting, is inevitably subject to noisy inputs arising from inaccurate 3D pose estimations and occlusions. To address these challenges, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Saeed Saadatnejad , Ali Rasekh , Mohammadreza Mofayezi , Yasamin Medghalchi , Sara Rajabzadeh , Taylor Mordan , Alexandre Alahi

Estimating robot pose and joint angles is significant in advanced robotics, enabling applications like robot collaboration and online hand-eye calibration.However, the introduction of unknown joint angles makes prediction more complex than…

Robotics · Computer Science 2024-03-28 Yang Tian , Jiyao Zhang , Guowei Huang , Bin Wang , Ping Wang , Jiangmiao Pang , Hao Dong

Rendering articulated objects while controlling their poses is critical to applications such as virtual reality or animation for movies. Manipulating the pose of an object, however, requires the understanding of its underlying structure,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Atsuhiro Noguchi , Umar Iqbal , Jonathan Tremblay , Tatsuya Harada , Orazio Gallo

Recently multi-view crowd counting using deep neural networks has been proposed to enable counting in large and wide scenes using multiple cameras. The current methods project the camera-view features to the average-height plane of the 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Qi Zhang , Antoni B. Chan

Monocular 3D human pose estimation remains a challenging and ill-posed problem, particularly in real-time settings and unconstrained environments. While direct imageto-3D approaches require large annotated datasets and heavy models,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Mohamed Adjel

Robust and accurate perception of humans in their 3D scene context is essential for integrating robots into everyday environments. Existing approaches, however, often fail to predict plausible and accurate human motion estimates that are…

Robotics · Computer Science 2026-05-26 Simon Schaefer , Joshua Näf , Stefan Leutenegger

In recent years, there has been rapid development in 3D generation models, opening up new possibilities for applications such as simulating the dynamic movements of 3D objects and customizing their behaviors. However, current 3D generative…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Fangfu Liu , Hanyang Wang , Shunyu Yao , Shengjun Zhang , Jie Zhou , Yueqi Duan