English
Related papers

Related papers: MC-VTON: Minimal Control Virtual Try-On Diffusion …

200 papers

Existing image-based virtual try-on (VTON) methods primarily focus on single-layer or multi-garment VTON, neglecting multi-layer VTON (ML-VTON), which involves dressing multiple layers of garments onto the human body with realistic…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Yang Yu , Yunze Deng , Yige Zhang , Yanjie Xiao , Youkun Ou , Wenhao Hu , Mingchao Li , Bin Feng , Wenyu Liu , Dandan Zheng , Jingdong Chen

Diffusion-based virtual try-on methods achieve photorealistic synthesis through cross-attention mechanisms that transfer garment features to target body regions. However, these approaches rely on implicit learning of spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Kosuke Takemoto , Takafumi Koshinaka

The Diffusion model has a strong ability to generate wild images. However, the model can just generate inaccurate images with the guidance of text, which makes it very challenging to directly apply the text-guided generative model for…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Shufang Zhang , Minxue Ni , Lei Wang , Wenxin Ding , Shuai Chen , Yuhong Liu

We present X-MDPT ($\underline{Cross}$-view $\underline{M}$asked $\underline{D}$iffusion $\underline{P}$rediction $\underline{T}$ransformers), a novel diffusion model designed for pose-guided human image generation. X-MDPT distinguishes…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Trung X. Pham , Zhang Kang , Chang D. Yoo

Diffusion Transformers (DiTs) have demonstrated exceptional capabilities in text-to-image synthesis. However, in the domain of controllable text-to-image generation using DiTs, most existing methods still rely on the ControlNet paradigm…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Shanyuan Liu , Jian Zhu , Junda Lu , Yue Gong , Liuzhuozheng Li , Bo Cheng , Yuhang Ma , Liebucha Wu , Xiaoyu Wu , Dawei Leng , Yuhui Yin

Virtual Try-On technology has garnered significant attention for its potential to transform the online fashion retail experience by allowing users to visualize how garments would look on them without physical trials. While recent advances…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Minh Tran , Johnmark Clements , Annie Prasanna , Tri Nguyen , Ngan Le

Virtual try-on focuses on adjusting the given clothes to fit a specific person seamlessly while avoiding any distortion of the patterns and textures of the garment. However, the clothing identity uncontrollability and training inefficiency…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Jiazheng Xing , Chao Xu , Yijie Qian , Yang Liu , Guang Dai , Baigui Sun , Yong Liu , Jingdong Wang

We present OOTDiffusion, a novel network architecture for realistic and controllable image-based virtual try-on (VTON). We leverage the power of pretrained latent diffusion models, designing an outfitting UNet to learn the garment detail…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Yuhao Xu , Tao Gu , Weifeng Chen , Chengcai Chen

Recent diffusion- and flow-based VTON methods achieve strong results with pretrained generative models, but their reliance on multi-step sampling incurs high inference cost, while existing acceleration methods largely overlook the intrinsic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Xianbing Sun , Jiahui Zhan , Liqing Zhang , Jianfu Zhang

The virtual try-on system has gained great attention due to its potential to give customers a realistic, personalized product presentation in virtualized settings. In this paper, we present PT-VTON, a novel pose-transfer-based framework for…

Computer Vision and Pattern Recognition · Computer Science 2021-11-25 Hanhan Zhou , Tian Lan , Guru Venkataramani

Virtual Try-On (VTON) has become a transformative technology, empowering users to experiment with fashion without ever having to physically try on clothing. However, existing methods often struggle with generating high-fidelity and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Ke Sun , Jian Cao , Qi Wang , Linrui Tian , Xindi Zhang , Lian Zhuo , Bang Zhang , Liefeng Bo , Wenbo Zhou , Weiming Zhang , Daiheng Gao

Diffusion Transformers (DiTs) with billions of model parameters form the backbone of popular image and video generation models like DALL.E, Stable-Diffusion and SORA. Though these models are necessary in many low-latency applications like…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Vignesh Sundaresha

Virtual try-on (VITON) aims to generate realistic images of a person wearing a target garment, requiring precise garment alignment in try-on regions and faithful preservation of identity and background in non-try-on regions. While latent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Junseo Park , Hyeryung Jang

With the development of Generative Adversarial Network, image-based virtual try-on methods have made great progress. However, limited work has explored the task of video-based virtual try-on while it is important in real-world applications.…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Xiaojing Zhong , Zhonghua Wu , Taizhe Tan , Guosheng Lin , Qingyao Wu

Video virtual try-on aims to transfer a clothing item onto the video of a target person. Directly applying the technique of image-based try-on to the video domain in a frame-wise manner will cause temporal-inconsistent outcomes while…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Zixun Fang , Wei Zhai , Aimin Su , Hongliang Song , Kai Zhu , Mao Wang , Yu Chen , Zhiheng Liu , Yang Cao , Zheng-Jun Zha

Diffusion models have become a leading approach for high-fidelity medical image synthesis. However, most existing methods for 3D medical image generation rely on convolutional U-Net backbones within latent diffusion frameworks. While…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Marvin Seyfarth , Salman Ul Hassan Dar , Yannik Frisch , Philipp Wild , Norbert Frey , Florian André , Sandy Engelhardt

Virtual try-on system under arbitrary human poses has huge application potential, yet raises quite a lot of challenges, e.g. self-occlusions, heavy misalignment among diverse poses, and diverse clothes textures. Existing methods aim at…

Computer Vision and Pattern Recognition · Computer Science 2019-03-01 Haoye Dong , Xiaodan Liang , Bochao Wang , Hanjiang Lai , Jia Zhu , Jian Yin

Diffusion Transformers (DiTs) introduce the transformer architecture to diffusion tasks for latent-space image generation. With an isotropic architecture that chains a series of transformer blocks, DiTs demonstrate competitive performance…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Yuchuan Tian , Zhijun Tu , Hanting Chen , Jie Hu , Chao Xu , Yunhe Wang

Video virtual try-on aims to generate realistic sequences that maintain garment identity and adapt to a person's pose and body shape in source videos. Traditional image-based methods, relying on warping and blending, struggle with complex…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Zijian He , Peixin Chen , Guangrun Wang , Guanbin Li , Philip H. S. Torr , Liang Lin

Image-based 3D Virtual Try-ON (VTON) aims to sculpt the 3D human according to person and clothes images, which is data-efficient (i.e., getting rid of expensive 3D data) but challenging. Recent text-to-3D methods achieve remarkable…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Zhenyu Xie , Haoye Dong , Yufei Gao , Zehua Ma , Xiaodan Liang