中文
相关论文

相关论文: CatV2TON: Taming Diffusion Transformers for Vision…

200 篇论文

The fashion industry is increasingly leveraging computer vision and deep learning technologies to enhance online shopping experiences and operational efficiencies. In this paper, we address the challenge of generating high-fidelity tiled…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Ioannis Xarchakos , Theodoros Koukopoulos

Virtual try-on is a promising application of computer graphics and human computer interaction that can have a profound real-world impact especially during this pandemic. Existing image-based works try to synthesize a try-on image from a…

图形学 · 计算机科学 2021-09-13 Toby Chong , I-Chao Shen , Nobuyuki Umetani , Takeo Igarashi

Image-based virtual try-on systems,which fit new garments onto human portraits,are gaining research attention.An ideal pipeline should preserve the static features of clothes(like textures and logos)while also generating dynamic…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Yanlong Zang , Han Yang , Jiaxu Miao , Yi Yang

Virtual try-on technology has become increasingly important in the fashion and retail industries, enabling the generation of high-fidelity garment images that adapt seamlessly to target human models. While existing methods have achieved…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Ming Meng , Qi Dong , Jiajie Li , Zhe Zhu , Xingyu Wang , Zhaoxin Fan , Wei Zhao , Wenjun Wu

Generative Adversarial Networks (GANs) dominate the research field in image-based virtual try-on, but have not resolved problems such as unnatural deformation of garments and the blurry generation quality. While the generative quality of…

计算机视觉与模式识别 · 计算机科学 2024-04-29 Jianhao Zeng , Dan Song , Weizhi Nie , Hongshuo Tian , Tongtong Wang , Anan Liu

Deep learning based virtual try-on system has achieved some encouraging progress recently, but there still remain several big challenges that need to be solved, such as trying on arbitrary clothes of all types, trying on the clothes from…

计算机视觉与模式识别 · 计算机科学 2021-11-25 Yu Liu , Mingbo Zhao , Zhao Zhang , Haijun Zhang , Shuicheng Yan

Multi-pose virtual try-on (MPVTON) aims to fit a target garment onto a person at a target pose. Compared to traditional virtual try-on (VTON) that fits the garment but keeps the pose unchanged, MPVTON provides a better try-on experience,…

计算机视觉与模式识别 · 计算机科学 2022-11-22 Sen He , Yi-Zhe Song , Tao Xiang

As online shopping continues to grow, the demand for Virtual Try-On (VTON) technology has surged, allowing customers to visualize products on themselves by overlaying product images onto their own photos. An essential yet challenging…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Qi Li , Shuwen Qiu , Julien Han , Xingzi Xu , Mehmet Saygin Seyfioglu , Kee Kiat Koo , Karim Bouyarmane

Virtual try-on can significantly improve the garment shopping experiences in both online and in-store scenarios, attracting broad interest in computer vision. However, to achieve high-fidelity try-on performance, most state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2024-02-06 Yunfang Niu , Dong Yi , Lingxiang Wu , Zhiwei Liu , Pengxiang Cai , Jinqiao Wang

Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealing alternative. Yet,…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Junpeng Jiang , Gangyi Hong , Miao Zhang , Hengtong Hu , Kun Zhan , Rui Shao , Liqiang Nie

Recent Diffusion Transformers (DiTs) have shown impressive capabilities in generating high-quality single-modality content, including images, videos, and audio. However, it is still under-explored whether the transformer-based diffuser can…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Kai Wang , Shijian Deng , Jing Shi , Dimitrios Hatzinakos , Yapeng Tian

We introduce RefTon, a flux-based person-to-person virtual try-on framework that enhances garment realism through unpaired visual references. Unlike conventional approaches that rely on complex auxiliary inputs such as body parsing and…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Liuzhuozheng Li , Yue Gong , Shanyuan Liu , Dengyang Jiang , Zanyi Wang , Bo Cheng , Yuhang Ma , Leibucha Wu , Dawei Leng , Yuhui Yin

Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Hanyue Lou , Jinxiu Liang , Minggui Teng , Yi Wang , Boxin Shi

Image-based virtual try-on aims to synthesize a naturally dressed person image with a clothing image, which revolutionizes online shopping and inspires related topics within image generation, showing both research significance and…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Dan Song , Xuanpu Zhang , Juan Zhou , Weizhi Nie , Ruofeng Tong , Mohan Kankanhalli , An-An Liu

Diffusion Transformers (DiTs) can generate short photorealistic videos, yet directly training and sampling longer videos with full attention across the video remains computationally challenging. Alternative methods break long videos down…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Bhishma Dedhia , David Bourgin , Krishna Kumar Singh , Yuheng Li , Yan Kang , Zhan Xu , Niraj K. Jha , Yuchen Liu

With the rapid development of e-commerce and digital fashion, image-based virtual try-on (VTON) has attracted increasing attention. However, existing VTON models often suffer from artifacts such as garment distortion and body inconsistency,…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Xinyi Wei , Sijing Wu , Zitong Xu , Yunhao Li , Huiyu Duan , Xiongkuo Min , Guangtao Zhai

We introduce DiffusionTrend for virtual fashion try-on, which forgoes the need for retraining diffusion models. Using advanced diffusion models, DiffusionTrend harnesses latent information rich in prior information to capture the nuances of…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Wengyi Zhan , Mingbao Lin , Shuicheng Yan , Rongrong Ji

We present xGen-VideoSyn-1, a text-to-video (T2V) generation model capable of producing realistic scenes from textual descriptions. Building on recent advancements, such as OpenAI's Sora, we explore the latent diffusion model (LDM)…

Studies of virtual try-on (VITON) have been shown their effectiveness in utilizing the generative neural network for virtually exploring fashion products, and some of recent researches of VITON attempted to synthesize human image wearing…

计算机视觉与模式识别 · 计算机科学 2022-05-11 Soonchan Park , Jinah Park

Text-to-image (T2I) diffusion models have revolutionized visual content creation, but extending these capabilities to text-to-video (T2V) generation remains a challenge, particularly in preserving temporal consistency. Existing methods that…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Dohun Lee , Bryan S Kim , Geon Yeong Park , Jong Chul Ye