中文
相关论文

相关论文: Single Stage Virtual Try-on via Deformable Attenti…

200 篇论文

Virtual Try-On is a promising research area with broad applications in e-commerce and everyday life, enabling users to visualize garments on themselves or others before purchase. Most existing methods depend on predefined or user-specified…

图形学 · 计算机科学 2026-04-01 Mengqi Zhang , Qi Li , Mehmet Saygin Seyfioglu , Karim Bouyarmane

Per-garment virtual try-on methods collect garment-specific datasets and train networks tailored to each garment to achieve superior results. However, these approaches often struggle with loose-fitting garments due to two key limitations:…

图形学 · 计算机科学 2025-09-05 Zaiqiang Wu , I-Chao Shen , Takeo Igarashi

Recent learning-based methods for event-based optical flow estimation utilize cost volumes for pixel matching but suffer from redundant computations and limited scalability to higher resolutions for flow refinement. In this work, we take…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Daikun Liu , Lei Cheng , Teng Wang , changyin Sun

Sequential multi-step cloth manipulation is a challenging problem in robotic manipulation, requiring a robot to perceive the cloth state and plan a sequence of chained actions leading to the desired state. Most previous works address this…

机器人学 · 计算机科学 2023-01-10 Kai Mo , Chongkun Xia , Xueqian Wang , Yuhong Deng , Xuehai Gao , Bin Liang

Diffusion-based virtual try-on methods achieve photorealistic synthesis through cross-attention mechanisms that transfer garment features to target body regions. However, these approaches rely on implicit learning of spatial…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Kosuke Takemoto , Takafumi Koshinaka

The 2D image-based virtual try-on has aroused increased interest from the multimedia and computer vision fields due to its enormous commercial value. Nevertheless, most existing image-based virtual try-on approaches directly combine the…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Bin Ren , Hao Tang , Fanyang Meng , Runwei Ding , Philip H. S. Torr , Nicu Sebe

Video virtual try-on aims to replace the clothing of a person in a video with a target garment. Current dual-branch architectures have achieved significant success in diffusion models based on the U-Net; however, adapting them to diffusion…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Yanjie Pan , Qingdong He , Lidong Wang , Bo Peng , Mingmin Chi

In this paper, we introduce the novel state-of-the-art Dual-attention Transformer and Discriminative Flow (DADF) framework for visual anomaly detection. Based on only normal knowledge, visual anomaly detection has wide applications in…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Haiming Yao , Wei Luo , Wenyong Yu

Latest advances have achieved realistic virtual try-on (VTON) through localized garment inpainting using latent diffusion models, significantly enhancing consumers' online shopping experience. However, existing VTON technologies neglect the…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Fei Shen , Xin Jiang , Xin He , Hu Ye , Cong Wang , Xiaoyu Du , Zechao Li , Jinhui Tang

We present a unified formulation and model for three motion and 3D perception tasks: optical flow, rectified stereo matching and unrectified stereo depth estimation from posed images. Unlike previous specialized architectures for each…

计算机视觉与模式识别 · 计算机科学 2023-07-27 Haofei Xu , Jing Zhang , Jianfei Cai , Hamid Rezatofighi , Fisher Yu , Dacheng Tao , Andreas Geiger

The virtual try-on task refers to fitting the clothes from one image onto another portrait image. In this paper, we focus on virtual accessory try-on, which fits accessory (e.g., glasses, ties) onto a face or portrait image. Unlike clothing…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Junhong Gou , Bo Zhang , Li Niu , Jianfu Zhang , Jianlou Si , Chen Qian , Liqing Zhang

Given two images depicting a person and a garment worn by another person, our goal is to generate a visualization of how the garment might look on the input person. A key challenge is to synthesize a photorealistic detail-preserving…

计算机视觉与模式识别 · 计算机科学 2023-06-16 Luyang Zhu , Dawei Yang , Tyler Zhu , Fitsum Reda , William Chan , Chitwan Saharia , Mohammad Norouzi , Ira Kemelmacher-Shlizerman

Deeply learned representations have achieved superior image retrieval performance in a retrieve-then-rerank manner. Recent state-of-the-art single stage model, which heuristically fuses local and global features, achieves promising…

计算机视觉与模式识别 · 计算机科学 2022-07-04 Yuxin Song , Ruolin Zhu , Min Yang , Dongliang He

Image-based virtual try-on is an increasingly important task for online shopping. It aims to synthesize images of a specific person wearing a specified garment. Diffusion model-based approaches have recently become popular, as they are…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Xu Yang , Changxing Ding , Zhibin Hong , Junhao Huang , Jin Tao , Xiangmin Xu

Performing facial expression transfer under one-shot setting has been increasing in popularity among research community with a focus on precise control of expressions. Existing techniques showcase compelling results in perceiving…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Siddharth Nijhawan , Takuya Yashima , Tamaki Kojima

Conventional physically based rendering (PBR) pipelines generate photorealistic images through computationally intensive light transport simulations. Although recent deep learning approaches leverage diffusion model priors with geometry…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Shenghao Zhang , Runtao Liu , Christopher Schroers , Yang Zhang

Pose-guided person image synthesis aims to synthesize person images by transforming reference images into target poses. In this paper, we observe that the commonly used spatial transformation blocks have complementary advantages. We propose…

计算机视觉与模式识别 · 计算机科学 2021-08-05 Yurui Ren , Yubo Wu , Thomas H. Li , Shan Liu , Ge Li

Virtual Try-On (VTON) has become a crucial tool in ecommerce, enabling the realistic simulation of garments on individuals while preserving their original appearance and pose. Early VTON methods relied on single generative networks, but…

计算机视觉与模式识别 · 计算机科学 2025-01-10 Shuliang Ning , Yipeng Qin , Xiaoguang Han

We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation within a unified autoregressive architecture, NextFlow…

Generating high-dimensional visual modalities is a computationally intensive task. A common solution is progressive generation, where the outputs are synthesized in a coarse-to-fine spectral autoregressive manner. While diffusion models…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Moayed Haji-Ali , Willi Menapace , Ivan Skorokhodov , Arpit Sahni , Sergey Tulyakov , Vicente Ordonez , Aliaksandr Siarohin