English

3D-C2FT: Coarse-to-fine Transformer for Multi-view 3D Reconstruction

Computer Vision and Pattern Recognition 2023-01-18 v2 Artificial Intelligence Machine Learning

Abstract

Recently, the transformer model has been successfully employed for the multi-view 3D reconstruction problem. However, challenges remain on designing an attention mechanism to explore the multiview features and exploit their relations for reinforcing the encoding-decoding modules. This paper proposes a new model, namely 3D coarse-to-fine transformer (3D-C2FT), by introducing a novel coarse-to-fine(C2F) attention mechanism for encoding multi-view features and rectifying defective 3D objects. C2F attention mechanism enables the model to learn multi-view information flow and synthesize 3D surface correction in a coarse to fine-grained manner. The proposed model is evaluated by ShapeNet and Multi-view Real-life datasets. Experimental results show that 3D-C2FT achieves notable results and outperforms several competing models on these datasets.

Keywords

Cite

@article{arxiv.2205.14575,
  title  = {3D-C2FT: Coarse-to-fine Transformer for Multi-view 3D Reconstruction},
  author = {Leslie Ching Ow Tiong and Dick Sigmund and Andrew Beng Jin Teoh},
  journal= {arXiv preprint arXiv:2205.14575},
  year   = {2023}
}

Comments

Accepted by Asian Conference on Computer Vision (ACCV) 2022

R2 v1 2026-06-24T11:32:07.688Z