English

MultiGO++: Monocular 3D Clothed Human Reconstruction via Geometry-Texture Collaboration

Computer Vision and Pattern Recognition 2026-03-06 v1

Abstract

Monocular 3D clothed human reconstruction aims to generate a complete and realistic textured 3D avatar from a single image. Existing methods are commonly trained under multi-view supervision with annotated geometric priors, and during inference, these priors are estimated by the pre-trained network from the monocular input. These methods are constrained by three key limitations: texturally by unavailability of training data, geometrically by inaccurate external priors, and systematically by biased single-modality supervision, all leading to suboptimal reconstruction. To address these issues, we propose a novel reconstruction framework, named MultiGO++, which achieves effective systematic geometry-texture collaboration. It consists of three core parts: (1) A multi-source texture synthesis strategy that constructs 15,000+ 3D textured human scans to improve the performance on texture quality estimation in challenge scenarios; (2) A region-aware shape extraction module that extracts and interacts features of each body region to obtain geometry information and a Fourier geometry encoder that mitigates the modality gap to achieve effective geometry learning; (3) A dual reconstruction U-Net that leverages geometry-texture collaborative features to refine and generate high-fidelity textured 3D human meshes. Extensive experiments on two benchmarks and many in-the-wild cases show the superiority of our method over state-of-the-art approaches.

Keywords

Cite

@article{arxiv.2603.04993,
  title  = {MultiGO++: Monocular 3D Clothed Human Reconstruction via Geometry-Texture Collaboration},
  author = {Nanjie Yao and Gangjian Zhang and Wenhao Shen and Jian Shu and Yu Feng and Hao Wang},
  journal= {arXiv preprint arXiv:2603.04993},
  year   = {2026}
}
R2 v1 2026-07-01T11:04:37.985Z