中文
相关论文

相关论文: CatV2TON: Taming Diffusion Transformers for Vision…

200 篇论文

Image-based virtual try-on aims to transfer an in-shop clothing image to a person image. Most existing methods adopt a single global deformation to perform clothing warping directly, which lacks fine-grained modeling of in-shop clothing and…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Shengping Zhang , Xiaoyu Han , Weigang Zhang , Xiangyuan Lan , Hongxun Yao , Qingming Huang

Video-to-video diffusion models achieve impressive single-turn editing performance, but practical editing workflows are inherently iterative. When edits are applied sequentially, existing models treat each turn independently, often causing…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Dohun Lee , Chun-Hao Paul Huang , Xuelin Chen , Jong Chul Ye , Duygu Ceylan , Hyeonho Jeong

We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification. Our model extracts spatio-temporal tokens from the input video, which are then encoded by a series of…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Anurag Arnab , Mostafa Dehghani , Georg Heigold , Chen Sun , Mario Lučić , Cordelia Schmid

Video virtual try-on aims to seamlessly replace the clothing of a person in a source video with a target garment. Despite significant progress in this field, existing approaches still struggle to maintain continuity and reproduce garment…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Jinjuan Wang , Wenzhang Sun , Ming Li , Yun Zheng , Fanyao Li , Zhulin Tao , Donglin Di , Hao Li , Wei Chen , Xianglin Huang

The development of virtual try-on has revolutionized online shopping by allowing customers to visualize themselves in various fashion items, thus extending the in-store try-on experience to the cyber space. Although virtual try-on has…

多媒体 · 计算机科学 2024-04-23 Mingzhe Yu , Yunshan Ma , Lei Wu , Kai Cheng , Xue Li , Lei Meng , Tat-Seng Chua

The growing digital landscape of fashion e-commerce calls for interactive and user-friendly interfaces for virtually trying on clothes. Traditional try-on methods grapple with challenges in adapting to diverse backgrounds, poses, and…

Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remains a persistent bottleneck. Traditional metrics struggle to quantify fine-grained texture…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Jin Li , Tao Chen , Shuai Jiang , Weijie Wang , Jingwen Luo , Chenhui Wu

Video virtual try-on technology provides a cost-effective solution for creating marketing videos in fashion e-commerce. However, its practical adoption is hindered by two critical limitations. First, the reliance on a single garment image…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Jianhao Zeng , Yancheng Bai , Ruidong Chen , Xuanpu Zhang , Lei Sun , Dongyang Jin , Ryan Xu , Nannan Zhang , Dan Song , Xiangxiang Chu

Although image-based virtual try-on has made considerable progress, emerging approaches still encounter challenges in producing high-fidelity and robust fitting images across diverse scenarios. These methods often struggle with issues such…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Boyuan Jiang , Xiaobin Hu , Donghao Luo , Qingdong He , Chengming Xu , Jinlong Peng , Jiangning Zhang , Chengjie Wang , Yunsheng Wu , Yanwei Fu

In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to explore extending this unification to videos. The…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Yuying Ge , Yizhuo Li , Yixiao Ge , Ying Shan

Despite remarkable progress in image-based virtual try-on systems, generating realistic and robust fitting images for cross-category virtual try-on remains a challenging task. The primary difficulty arises from the absence of human-like…

计算机视觉与模式识别 · 计算机科学 2025-02-21 Donghao Luo , Yujie Liang , Xu Peng , Xiaobin Hu , Boyuan Jiang , Chengming Xu , Taisong Jin , Chengjie Wang , Yanwei Fu

Image-based virtual try-on aims to fit an in-shop garment onto a clothed person image. Garment warping, which aligns the target garment with the corresponding body parts in the person image, is a crucial step in achieving this goal.…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Sanhita Pathak , Vinay Kaushik , Brejesh Lall

Existing image-based virtual try-on methods directly transfer specific clothing to a human image without utilizing clothing attributes to refine the transferred clothing geometry and textures, which causes incomplete and blurred clothing…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Xiaoyu Han , Shengping Zhang , Qinglin Liu , Zonglin Li , Chenyang Wang

Generating a virtual try-on image from in-shop clothing images and a model person's snapshot is a challenging task because the human body and clothes have high flexibility in their shapes. In this paper, we develop a Virtual Try-on…

计算机视觉与模式识别 · 计算机科学 2019-11-20 Shion Honda

Video face swapping is becoming increasingly popular across various applications, yet existing methods primarily focus on static images and struggle with video face swapping because of temporal consistency and complex scenarios. In this…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Hao Shao , Shulun Wang , Yang Zhou , Guanglu Song , Dailan He , Shuo Qin , Zhuofan Zong , Bingqi Ma , Yu Liu , Hongsheng Li

Realistic virtual try-on (VTON) concerns not only faithful rendering of garment details but also coordination of the style. Prior art typically pursues the former, but neglects a key factor that shapes the holistic style -- garment fit.…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Lu Yang , Yicheng Liu , Yanan Li , Xiang Bai , Hao Lu

Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation. However, generating smooth and rational transition videos…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Zuhao Yang , Jiahui Zhang , Yingchen Yu , Shijian Lu , Song Bai

Virtual try-on aims to generate a photo-realistic fitting result given an in-shop garment and a reference person image. Existing methods usually build up multi-stage frameworks to deal with clothes warping and body blending respectively, or…

计算机视觉与模式识别 · 计算机科学 2022-07-20 Shuai Bai , Huiling Zhou , Zhikang Li , Chang Zhou , Hongxia Yang

Diffusion-based \textit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional conditions, into a temporally coherent video. Compared with broader video generation…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Xianlong Wang , Wenbo Pan , Shijia Zhou , Ke Li , Yuqi Wang , Zeyu Ye , Hangtao Zhang , Leo Yu Zhang , Xiaohua Jia

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Linfeng Tang , Yeda Wang , Meiqi Gong , Zizhuo Li , Yuxin Deng , Xunpeng Yi , Chunyu Li , Han Xu , Hao Zhang , Jiayi Ma