English

MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction

Computer Vision and Pattern Recognition 2026-03-20 v1

Abstract

Reconstructing articulated 3D objects from a single image requires jointly inferring object geometry, part structure, and motion parameters from limited visual evidence. A key difficulty lies in the entanglement between motion cues and object structure, which makes direct articulation regression unstable. Existing methods address this challenge through multi-view supervision, retrieval-based assembly, or auxiliary video generation, often sacrificing scalability or efficiency. We present MonoArt, a unified framework grounded in progressive structural reasoning. Rather than predicting articulation directly from image features, MonoArt progressively transforms visual observations into canonical geometry, structured part representations, and motion-aware embeddings within a single architecture. This structured reasoning process enables stable and interpretable articulation inference without external motion templates or multi-stage pipelines. Extensive experiments on PartNet-Mobility demonstrate that OM achieves state-of-the-art performance in both reconstruction accuracy and inference speed. The framework further generalizes to robotic manipulation and articulated scene reconstruction.

Keywords

Cite

@article{arxiv.2603.19231,
  title  = {MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction},
  author = {Haitian Li and Haozhe Xie and Junxiang Xu and Beichen Wen and Fangzhou Hong and Ziwei Liu},
  journal= {arXiv preprint arXiv:2603.19231},
  year   = {2026}
}

Comments

Project page: https://lihaitian.com/MonoArt

R2 v1 2026-07-01T11:28:40.348Z