English

Axolotl3D: a Unified Framework for Faithful 3D Shape Completion

Computer Vision and Pattern Recognition 2026-07-22 v1

Abstract

Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and diffusion architectures. However, they assume complete visibility and single-view inputs, limiting applicability in multi-view, occluded, or editing scenarios. Although prior works address these challenges individually, they lack a unified framework for controllable 3D completion under diverse conditioning signals. We present Axolotl3D, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud. The point cloud serves as a geometric anchor promoting faithful shape completion, while camera parameters ensure consistent multi-view alignment in a shared 3D coordinate system. A unified training strategy synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning. Experiments on Toys4K and OmniObject3D demonstrate state-of-the-art performance under both clean and occluded settings, as well as strong results in real-world reconstruction and geometry-consistent editing.

Cite

@article{arxiv.2607.20660,
  title  = {Axolotl3D: a Unified Framework for Faithful 3D Shape Completion},
  author = {Anita Hu and Maria Shugrina},
  journal= {arXiv preprint arXiv:2607.20660},
  year   = {2026}
}

Comments

Accepted to ECCV 2026