中文
相关论文

相关论文: CatV2TON: Taming Diffusion Transformers for Vision…

200 篇论文

Given two images depicting a person and a garment worn by another person, our goal is to generate a visualization of how the garment might look on the input person. A key challenge is to synthesize a photorealistic detail-preserving…

计算机视觉与模式识别 · 计算机科学 2023-06-16 Luyang Zhu , Dawei Yang , Tyler Zhu , Fitsum Reda , William Chan , Chitwan Saharia , Mohammad Norouzi , Ira Kemelmacher-Shlizerman

A virtual try-on method takes a product image and an image of a model and produces an image of the model wearing the product. Most methods essentially compute warps from the product image to the model image and combine using image…

计算机视觉与模式识别 · 计算机科学 2020-03-30 Kedan Li , Min Jin Chong , Jingen Liu , David Forsyth

Vision Transformers achieve impressive accuracy across a range of visual recognition tasks. Unfortunately, their accuracy frequently comes with high computational costs. This is a particular issue in video recognition, where models are…

计算机视觉与模式识别 · 计算机科学 2023-08-28 Matthew Dutson , Yin Li , Mohit Gupta

Diffusion models have transformed the image-to-image (I2I) synthesis and are now permeating into videos. However, the advancement of video-to-video (V2V) synthesis has been hampered by the challenge of maintaining temporal consistency…

计算机视觉与模式识别 · 计算机科学 2024-01-01 Feng Liang , Bichen Wu , Jialiang Wang , Licheng Yu , Kunpeng Li , Yinan Zhao , Ishan Misra , Jia-Bin Huang , Peizhao Zhang , Peter Vajda , Diana Marculescu

We present Vchitect-2.0, a parallel transformer architecture designed to scale up video diffusion models for large-scale text-to-video generation. The overall Vchitect-2.0 system has several key designs. (1) By introducing a novel…

Text-to-video (T2V) generation technology holds potential to transform multiple domains such as education, marketing, entertainment, and assistive technologies for individuals with visual or reading comprehension challenges, by creating…

图形学 · 计算机科学 2025-10-07 Nilay Kumar , Priyansh Bhandari , G. Maragatham

Modern text-to-video synthesis models demonstrate coherent, photorealistic generation of complex videos from a text description. However, most existing models lack fine-grained control over camera movement, which is critical for downstream…

View transformation robustness (VTR) is critical for deep-learning-based multi-view 3D object reconstruction models, which indicates the methods' stability under inputs with various view transformations. However, existing research seldom…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Qi Zhang , Zhouhang Luo , Tao Yu , Hui Huang

Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constraints that remain challenging for perspective video generators due to their limited field of…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Ting-Hsuan Chen , Ying-Huan Chen , Tao Tu , Jie-Ying Lee , Cho-Ying Wu , Fangzhou Lin , Hengyuan Zhang , David Paz , Xinyu Huang , Yuliang Guo , Yu-Lun Liu , Yue Wang , Liu Ren

Although diffusion-based image virtual try-on has made considerable progress, emerging approaches still struggle to effectively address the issue of hand occlusion (i.e., clothing regions occluded by the hand part), leading to a notable…

计算机视觉与模式识别 · 计算机科学 2024-08-28 Yujie Liang , Xiaobin Hu , Boyuan Jiang , Donghao Luo , Kai WU , Wenhui Han , Taisong Jin , Chengjie Wang

This paper introduces MMTryon, a multi-modal multi-reference VIrtual Try-ON (VITON) framework, which can generate high-quality compositional try-on results by taking a text instruction and multiple garment images as inputs. Our MMTryon…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Xujie Zhang , Ente Lin , Xiu Li , Yuxuan Luo , Michael Kampffmeyer , Xin Dong , Xiaodan Liang

Image-based virtual try-on is challenging in fitting a target in-shop clothes into a reference person under diverse human poses. Previous works focus on preserving clothing details ( e.g., texture, logos, patterns ) when transferring…

计算机视觉与模式识别 · 计算机科学 2022-10-18 Bingwen Hu , Ping Liu , Zhedong Zheng , Mingwu Ren

Given a person and a garment image, virtual try-on (VTO) aims to synthesize a realistic image of the person wearing the garment, while preserving their original pose and identity. Although recent VTO methods excel at visualizing garment…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Johanna Karras , Yuanhao Wang , Yingwei Li , Ira Kemelmacher-Shlizerman

While diffusion-based methods have shown impressive capabilities in capturing diverse and complex hairstyles, their ability to generate consistent and high-quality multi-view outputs -- crucial for real-world applications such as digital…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Kuiyuan Sun , Yuxuan Zhang , Jichao Zhang , Jiaming Liu , Wei Wang , Niculae Sebe , Yao Zhao

Deep video models, for example, 3D CNNs or video transformers, have achieved promising performance on sparse video tasks, i.e., predicting one result per video. However, challenges arise when adapting existing deep video models to dense…

计算机视觉与模式识别 · 计算机科学 2024-02-15 Guanxiong Sun , Yang Hua , Guosheng Hu , Neil Robertson

Learning efficient and expressive visual representation has long been the pursuit of computer vision research. While Vision Transformers (ViTs) gradually replace traditional Convolutional Neural Networks (CNNs) as more scalable vision…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Quan Kong , Yanru Xiao , Yuhao Shen , Cong Wang

Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Xingrui Wang , Xin Li , Yaosi Hu , Hanxin Zhu , Chen Hou , Cuiling Lan , Zhibo Chen

In this study, we explore Transformer-based diffusion models for image and video generation. Despite the dominance of Transformer architectures in various fields due to their flexibility and scalability, the visual generative domain…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Shoufa Chen , Mengmeng Xu , Jiawei Ren , Yuren Cong , Sen He , Yanping Xie , Animesh Sinha , Ping Luo , Tao Xiang , Juan-Manuel Perez-Rua

In this paper, we first investigate a visual quality degradation problem observed in recent high-resolution virtual try-on approach. The tendency is empirically found that the textures of clothes are squeezed at the sleeve, as visualized in…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Sang-Heon Shim , Jiwoo Chung , Jae-Pil Heo

Advances in diffusion-based video generation models, while significantly improving human animation, poses threats of misuse through the creation of fake videos from a specific person's photo and text prompts. Recent efforts have focused on…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Duc Vu , Anh Nguyen , Chi Tran , Anh Tran