English

CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling

Computer Vision and Pattern Recognition 2025-11-25 v2

Abstract

We introduce Cupid, a generative 3D reconstruction framework that jointly models the full distribution over both canonical objects and camera poses. Our two-stage flow-based model first generates a coarse 3D structure and 2D-3D correspondences to estimate the camera pose robustly. Conditioned on this pose, a refinement stage injects pixel-aligned image features directly into the generative process, marrying the rich prior of a generative model with the geometric fidelity of reconstruction. This strategy achieves exceptional faithfulness, outperforming state-of-the-art reconstruction methods by over 3 dB PSNR and 10% in Chamfer Distance. As a unified generative model that decouples the object and camera pose, Cupid naturally extends to multi-view and scene-level reconstruction tasks without requiring post-hoc optimization or fine-tuning.

Keywords

Cite

@article{arxiv.2510.20776,
  title  = {CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling},
  author = {Binbin Huang and Haobin Duan and Yiqun Zhao and Zibo Zhao and Yi Ma and Shenghua Gao},
  journal= {arXiv preprint arXiv:2510.20776},
  year   = {2025}
}

Comments

project page at https://cupid3d.github.io

R2 v1 2026-07-01T07:02:35.351Z