English
Related papers

Related papers: GEOPARD: Geometric Pretraining for Articulation Pr…

200 papers

3D models of manufactured objects are important for populating virtual worlds and for synthetic data generation for vision and robotics. To be most useful, such objects should be articulated: their parts should move when interacted with.…

Graphics · Computer Science 2022-06-20 Xianghao Xu , Yifan Ruan , Srinath Sridhar , Daniel Ritchie

Articulated 3D object generation is fundamental for creating realistic, functional, and interactable virtual assets which are not simply static. We introduce MeshArt, a hierarchical transformer-based approach to generate articulated 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Daoyi Gao , Yawar Siddiqui , Lei Li , Angela Dai

We address the task of simultaneous part-level reconstruction and motion parameter estimation for articulated objects. Given two sets of multi-view images of an object in two static articulation states, we decouple the movable part from the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Jiayi Liu , Ali Mahdavi-Amiri , Manolis Savva

Our work aims to reconstruct hand-held objects given a single RGB image. In contrast to prior works that typically assume known 3D templates and reduce the problem to 3D pose estimation, our work reconstructs generic hand-held object…

Computer Vision and Pattern Recognition · Computer Science 2022-04-15 Yufei Ye , Abhinav Gupta , Shubham Tulsiani

We propose a transformer-based neural network architecture for multi-object 3D reconstruction from RGB videos. It relies on two alternative ways to represent its knowledge: as a global 3D grid of features and an array of view-specific 2D…

Computer Vision and Pattern Recognition · Computer Science 2022-08-29 Michał J. Tyszkiewicz , Kevis-Kokitsi Maninis , Stefan Popov , Vittorio Ferrari

We introduce ART, Articulated Reconstruction Transformer -- a category-agnostic, feed-forward model that reconstructs complete 3D articulated objects from only sparse, multi-state RGB images. Previous methods for articulated object…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Zizhang Li , Cheng Zhang , Zhengqin Li , Henry Howard-Jenkins , Zhaoyang Lv , Chen Geng , Jiajun Wu , Richard Newcombe , Jakob Engel , Zhao Dong

Reconstructing articulated objects is essential for building digital twins of interactive environments. However, prior methods typically decouple geometry and motion by first reconstructing object shape in distinct states and then…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Licheng Shen , Saining Zhang , Honghan Li , Peilin Yang , Zihao Huang , Zongzheng Zhang , Hao Zhao

Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require precise 3D reasoning. We propose GeoPredict, a geometry-aware…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Jingjing Qian , Boyao Han , Chen Shi , Lei Xiao , Long Yang , Shaoshuai Shi , Li Jiang

Understanding and manipulating articulated objects, such as doors and drawers, is crucial for robots operating in human environments. We wish to develop a system that can learn to articulate novel objects with no prior interaction, after…

Robotics · Computer Science 2024-05-03 Harry Zhang , Ben Eisner , David Held

This paper focuses on the problem of learning 6-DOF grasping with a parallel jaw gripper in simulation. We propose the notion of a geometry-aware representation in grasping based on the assumption that knowledge of 3D geometry is at the…

Precisely grasping and reconstructing articulated objects is key to enabling general robotic manipulation. In this paper, we propose CenterArt, a novel approach for simultaneous 3D shape reconstruction and 6-DoF grasp estimation of…

Robotics · Computer Science 2024-04-24 Sassan Mokhtar , Eugenio Chisari , Nick Heppert , Abhinav Valada

Monocular 3D human pose estimation technologies have the potential to greatly increase the availability of human movement data. The best-performing models for single-image 2D-3D lifting use graph convolutional networks (GCNs) that typically…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Sebastian Lutz , Richard Blythman , Koustav Ghosal , Matthew Moynihan , Ciaran Simms , Aljosa Smolic

We present GTT-Net, a supervised learning framework for the reconstruction of sparse dynamic 3D geometry. We build on a graph-theoretic formulation of the generalized trajectory triangulation problem, where non-concurrent multi-view imaging…

Computer Vision and Pattern Recognition · Computer Science 2021-09-09 Xiangyu Xu , Enrique Dunn

We present THUNDR, a transformer-based deep neural network methodology to reconstruct the 3d pose and shape of people, given monocular RGB images. Key to our methodology is an intermediate 3d marker representation, where we aim to combine…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Mihai Zanfir , Andrei Zanfir , Eduard Gabriel Bazavan , William T. Freeman , Rahul Sukthankar , Cristian Sminchisescu

We consider the problem of predicting the 3D shape, articulation, viewpoint, texture, and lighting of an articulated animal like a horse given a single test image as input. We present a new method, dubbed MagicPony, that learns this…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Shangzhe Wu , Ruining Li , Tomas Jakab , Christian Rupprecht , Andrea Vedaldi

We propose ArtiLatent, a generative framework that synthesizes human-made 3D objects with fine-grained geometry, accurate articulation, and realistic appearance. Our approach jointly models part geometry and articulation dynamics by…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Honghua Chen , Yushi Lan , Yongwei Chen , Xingang Pan

Recent feed-forward networks have achieved remarkable progress in sparse-view 3D reconstruction by predicting dense point maps directly from RGB images. However, they often suffer from geometric inconsistencies and limited fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yutong Chen , Yiming Wang , Xucong Zhang , Sergey Prokudin , Siyu Tang

The Geometric Algebra Transformer (GATr) is a versatile architecture for geometric deep learning based on projective geometric algebra. We generalize this architecture into a blueprint that allows one to construct a scalable transformer…

Machine Learning · Computer Science 2024-03-15 Pim de Haan , Taco Cohen , Johann Brehmer

Markerless motion capture algorithms require a 3D body with properly personalized skeleton dimension and/or body shape and appearance to successfully track a person. Unfortunately, many tracking methods consider model personalization a…

Computer Vision and Pattern Recognition · Computer Science 2016-10-24 Helge Rhodin , Nadia Robertini , Dan Casas , Christian Richardt , Hans-Peter Seidel , Christian Theobalt

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly,…