English

TexDreamer: Towards Zero-Shot High-Fidelity 3D Human Texture Generation

Computer Vision and Pattern Recognition 2024-03-20 v1

Abstract

Texturing 3D humans with semantic UV maps remains a challenge due to the difficulty of acquiring reasonably unfolded UV. Despite recent text-to-3D advancements in supervising multi-view renderings using large text-to-image (T2I) models, issues persist with generation speed, text consistency, and texture quality, resulting in data scarcity among existing datasets. We present TexDreamer, the first zero-shot multimodal high-fidelity 3D human texture generation model. Utilizing an efficient texture adaptation finetuning strategy, we adapt large T2I model to a semantic UV structure while preserving its original generalization capability. Leveraging a novel feature translator module, the trained model is capable of generating high-fidelity 3D human textures from either text or image within seconds. Furthermore, we introduce ArTicuLated humAn textureS (ATLAS), the largest high-resolution (1024 X 1024) 3D human texture dataset which contains 50k high-fidelity textures with text descriptions.

Keywords

Cite

@article{arxiv.2403.12906,
  title  = {TexDreamer: Towards Zero-Shot High-Fidelity 3D Human Texture Generation},
  author = {Yufei Liu and Junwei Zhu and Junshu Tang and Shijie Zhang and Jiangning Zhang and Weijian Cao and Chengjie Wang and Yunsheng Wu and Dongjin Huang},
  journal= {arXiv preprint arXiv:2403.12906},
  year   = {2024}
}

Comments

Project Page: https://ggxxii.github.io/texdreamer/

R2 v1 2026-06-28T15:26:01.852Z