English

TELA: Text to Layer-wise 3D Clothed Human Generation

Computer Vision and Pattern Recognition 2024-04-26 v1

Abstract

This paper addresses the task of 3D clothed human generation from textural descriptions. Previous works usually encode the human body and clothes as a holistic model and generate the whole model in a single-stage optimization, which makes them struggle for clothing editing and meanwhile lose fine-grained control over the whole generation process. To solve this, we propose a layer-wise clothed human representation combined with a progressive optimization strategy, which produces clothing-disentangled 3D human models while providing control capacity for the generation process. The basic idea is progressively generating a minimal-clothed human body and layer-wise clothes. During clothing generation, a novel stratified compositional rendering method is proposed to fuse multi-layer human models, and a new loss function is utilized to help decouple the clothing model from the human body. The proposed method achieves high-quality disentanglement, which thereby provides an effective way for 3D garment generation. Extensive experiments demonstrate that our approach achieves state-of-the-art 3D clothed human generation while also supporting cloth editing applications such as virtual try-on. Project page: http://jtdong.com/tela_layer/

Keywords

Cite

@article{arxiv.2404.16748,
  title  = {TELA: Text to Layer-wise 3D Clothed Human Generation},
  author = {Junting Dong and Qi Fang and Zehuan Huang and Xudong Xu and Jingbo Wang and Sida Peng and Bo Dai},
  journal= {arXiv preprint arXiv:2404.16748},
  year   = {2024}
}
R2 v1 2026-06-28T16:06:36.053Z