English
Related papers

Related papers: MMTryon: Multi-Modal Multi-Reference Control for H…

200 papers

Designing real and virtual garments is becoming extremely demanding with rapidly changing fashion trends and increasing need for synthesizing realistic dressed digital humans for various applications. This necessitates creating simple and…

Graphics · Computer Science 2018-07-17 Tuanfeng Y. Wang , Duygu Ceylan , Jovan Popovic , Niloy J. Mitra

Virtual Try-Off (VTOFF) is a challenging multimodal image generation task that aims to synthesize high-fidelity flat-lay garments under complex geometric deformation and rich high-frequency textures. Existing methods often rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Yihan Zhu , Mengying Ge

Recent virtual try-on approaches have advanced by finetuning pre-trained text-to-image diffusion models to leverage their powerful generative ability. However, the use of text prompts in virtual try-on remains underexplored. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Jeongho Kim , Hoiyeong Jin , Sunghyun Park , Jaegul Choo

Large Multimodal Models (LMMs) have recently demonstrated remarkable visual understanding performance on both vision-language and vision-centric tasks. However, they often fall short in integrating advanced, task-specific capabilities for…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Yufei Zhan , Hongyin Zhao , Yousong Zhu , Shurong Zheng , Fan Yang , Ming Tang , Jinqiao Wang

With the development of deep learning technology, virtual try-on technology has devel-oped important application value in the fields of e-commerce, fashion, and entertainment. The recently proposed Leffa technology has addressed the texture…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Sehyun Kim , Hye Jun Lee , Jiwoo Lee , Taemin Lee

Multimodal deep learning, especially vision-language models, have gained significant traction in recent years, greatly improving performance on many downstream tasks, including content moderation and violence detection. However, standard…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Zhuokai Zhao , Harish Palani , Tianyi Liu , Lena Evans , Ruth Toner

Large-scale Vision-and-Language (V+L) pre-training for representation learning has proven to be effective in boosting various downstream V+L tasks. However, when it comes to the fashion domain, existing V+L methods are inadequate as they…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Xiao Han , Licheng Yu , Xiatian Zhu , Li Zhang , Yi-Zhe Song , Tao Xiang

As virtual try-on (VTON) continues to advance, a growing number of real-world scenarios have emerged, pushing beyond the ability of the existing specialized VTON models. Meanwhile, universal multi-reference image editing models have…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Xiaoye Liang , Zhiyuan Qu , Mingye Zou , Jiaxin Liu , Lai Jiang , Mai Xu , Yiheng Zhu

The growing digital landscape of fashion e-commerce calls for interactive and user-friendly interfaces for virtually trying on clothes. Traditional try-on methods grapple with challenges in adapting to diverse backgrounds, poses, and…

Human-Computer Interaction · Computer Science 2024-02-06 Justin Blalock , David Munechika , Harsha Karanth , Alec Helbling , Pratham Mehta , Seongmin Lee , Duen Horng Chau

Recent advances in garment-centric image generation from text and image prompts based on diffusion models are impressive. However, existing methods lack support for various combinations of attire, and struggle to preserve the garment…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Xinghui Li , Qichao Sun , Pengze Zhang , Fulong Ye , Zhichao Liao , Wanquan Feng , Songtao Zhao , Qian He

Fashion video generation aims to synthesize temporally consistent videos from reference images of a designated character. Despite significant progress, existing diffusion-based methods only support a single reference image as input,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Xianghao Kong , Qiaosong Qi , Yuanbin Wang , Biaolong Chen , Aixi Zhang , Anyi Rao

We present CM3Leon (pronounced "Chameleon"), a retrieval-augmented, token-based, decoder-only multi-modal language model capable of generating and infilling both text and images. CM3Leon uses the CM3 multi-modal architecture but…

Industrial recommendation systems typically involve multiple scenarios, yet existing cross-domain (CDR) and multi-scenario (MSR) methods often require prohibitive resources and strict input alignment, limiting their extensibility. We…

Information Retrieval · Computer Science 2026-02-16 Xin Song , Zhilin Guan , Ruidong Han , Binghao Tang , Tianwen Chen , Bing Li , Zihao Li , Han Zhang , Fei Jiang , Qing Wang , Zikang Xu , Fengyi Li , Chunzhen Jing , Lei Yu , Wei Lin

Virtual try-on can significantly improve the garment shopping experiences in both online and in-store scenarios, attracting broad interest in computer vision. However, to achieve high-fidelity try-on performance, most state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Yunfang Niu , Dong Yi , Lingxiang Wu , Zhiwei Liu , Pengxiang Cai , Jinqiao Wang

This paper proposes a novel garment transfer method supervised with knowledge distillation from virtual try-on. Our method first reasons the transfer parsing to provide shape prior to downstream tasks. We employ a multi-phase teaching…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Naiyu Fang , Lemiao Qiu , Shuyou Zhang , Zili Wang , Kerui Hu , Jianrong Tan

Virtual try-on has made significant progress in recent years. This paper addresses how to achieve multifunctional virtual try-on guided solely by text instructions, including full outfit change and local editing. Previous methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Yujie Hu , Xuanyu Zhang , Weiqi Li , Jian Zhang

Our paper seeks to transfer the hairstyle of a reference image to an input photo for virtual hair try-on. We target a variety of challenges scenarios, such as transforming a long hairstyle with bangs to a pixie cut, which requires removing…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Sasikarn Khwanmuang , Pakkapon Phongthawee , Patsorn Sangkloy , Supasorn Suwajanakorn

In recent years, the fashion industry has increasingly adopted AI technologies to enhance customer experience, driven by the proliferation of e-commerce platforms and virtual applications. Among the various tasks, virtual try-on and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Fulvio Sanguigni , Davide Morelli , Marcella Cornia , Rita Cucchiara

Fashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay between clothing and the human body. In the context of fashion…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Alberto Baldrati , Davide Morelli , Marcella Cornia , Marco Bertini , Rita Cucchiara

Generative models have achieved impressive fidelity in text-to-image synthesis, yet struggle with complex compositional prompts involving multiple constraints. We introduce \textbf{M3 (Multi-Modal, Multi-Agent, Multi-Round)}, a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Bangji Yang , Ruihan Guo , Jiajun Fan , Chaoran Cheng , Ge Liu