English

To The Point: Correspondence-driven monocular 3D category reconstruction

Computer Vision and Pattern Recognition 2021-06-11 v1

Abstract

We present To The Point (TTP), a method for reconstructing 3D objects from a single image using 2D to 3D correspondences learned from weak supervision. We recover a 3D shape from a 2D image by first regressing the 2D positions corresponding to the 3D template vertices and then jointly estimating a rigid camera transform and non-rigid template deformation that optimally explain the 2D positions through the 3D shape projection. By relying on 3D-2D correspondences we use a simple per-sample optimization problem to replace CNN-based regression of camera pose and non-rigid deformation and thereby obtain substantially more accurate 3D reconstructions. We treat this optimization as a differentiable layer and train the whole system in an end-to-end manner. We report systematic quantitative improvements on multiple categories and provide qualitative results comprising diverse shape, pose and texture prediction examples. Project website: https://fkokkinos.github.io/to_the_point/.

Keywords

Cite

@article{arxiv.2106.05662,
  title  = {To The Point: Correspondence-driven monocular 3D category reconstruction},
  author = {Filippos Kokkinos and Iasonas Kokkinos},
  journal= {arXiv preprint arXiv:2106.05662},
  year   = {2021}
}
R2 v1 2026-06-24T03:03:08.636Z