Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplored. We introduce VLG, a vision-language-garment model that synthesizes garments from textual descriptions and visual imagery. Our experiments assess VLG's zero-shot generalization, investigating its ability to transfer web-scale reasoning to unseen garment styles and prompts. Preliminary results indicate promising transfer capabilities, highlighting the potential for multimodal foundation models to adapt effectively to specialized domains like fashion design.
@article{arxiv.2506.05210,
title = {Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation},
author = {Jan Ackermann and Kiyohiro Nakayama and Guandao Yang and Tong Wu and Gordon Wetzstein},
journal= {arXiv preprint arXiv:2506.05210},
year = {2025}
}
Comments
Presented at MMFM CVPRW'25, Project Page: https://www.computationalimaging.org/publications/vision-language-garment-models/