English
Related papers

Related papers: HF-VTON: High-Fidelity Virtual Try-On via Consiste…

200 papers

Visual Grounding (VG) aims to utilize given natural language queries to locate specific target objects within images. While current transformer-based approaches demonstrate strong localization performance in standard scene (i.e, scenarios…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Jiangnan Xie , Xiaolong Zheng , Liang Zheng

We introduce Spatial-Temporal Memory Networks for video object detection. At its core, a novel Spatial-Temporal Memory module (STMM) serves as the recurrent computation unit to model long-term temporal appearance and motion dynamics. The…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Fanyi Xiao , Yong Jae Lee

Virtual try-on aims to generate a photo-realistic fitting result given an in-shop garment and a reference person image. Existing methods usually build up multi-stage frameworks to deal with clothes warping and body blending respectively, or…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Shuai Bai , Huiling Zhou , Zhikang Li , Chang Zhou , Hongxia Yang

Virtual Try-ON (VTON) is a practical and widely-applied task, for which most of existing works focus on clothes. This paper presents OmniTry, a unified framework that extends VTON beyond garment to encompass any wearable objects, e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Yutong Feng , Linlin Zhang , Hengyuan Cao , Yiming Chen , Xiaoduan Feng , Jian Cao , Yuxiong Wu , Bin Wang

Image-based virtual try-on aims to synthesize a naturally dressed person image with a clothing image, which revolutionizes online shopping and inspires related topics within image generation, showing both research significance and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Dan Song , Xuanpu Zhang , Juan Zhou , Weizhi Nie , Ruofeng Tong , Mohan Kankanhalli , An-An Liu

Unpaired image translation algorithms can be used for sim2real tasks, but many fail to generate temporally consistent results. We present a new approach that combines differentiable rendering with image translation to achieve temporal…

Computer Vision and Pattern Recognition · Computer Science 2022-12-16 Ryan Burgert , Jinghuan Shang , Xiang Li , Michael Ryoo

Existing methods for Virtual Try-On (VTON) often struggle to preserve fine garment details, especially in unpaired settings where accurate person-garment correspondence is required. These methods do not explicitly enforce person-garment…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Jiyoung Kim , Youngjin Shin , Siyoon Jin , Dahyun Chung , Jisu Nam , Tongmin Kim , Jongjae Park , Hyeonwoo Kang , Seungryong Kim

Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision…

Computation and Language · Computer Science 2024-11-15 Xiang Zhang , Senyu Li , Ning Shi , Bradley Hauer , Zijun Wu , Grzegorz Kondrak , Muhammad Abdul-Mageed , Laks V. S. Lakshmanan

With the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visual prompt tuning is introduced as a parameter-efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Runjia Zeng , Cheng Han , Qifan Wang , Chunshu Wu , Tong Geng , Lifu Huang , Ying Nian Wu , Dongfang Liu

With the rapid advancements in Artificial Intelligence Generated Image (AGI) technology, the accurate assessment of their quality has become an increasingly vital requirement. Prevailing methods typically rely on cross-modal models like…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Qiang Li , Qingsen Yan , Haojian Huang , Peng Wu , Haokui Zhang , Yanning Zhang

Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Daixun Li , Weiying Xie , Mingxiang Cao , Yunke Wang , Yusi Zhang , Leyuan Fang , Yunsong Li , Chang Xu

This paper addresses a new virtual try-on problem of fitting any size of clothes to a reference person in the image domain. While previous image-based virtual try-on methods can produce highly natural try-on images, these methods fit the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Yohei Yamashita , Chihiro Nakatani , Norimichi Ukita

This paper presents a robust approach for a visual parallel tracking and mapping (PTAM) system that excels in challenging environments. Our proposed method combines the strengths of heterogeneous multi-modal visual sensors, including stereo…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Abanob Soliman , Fabien Bonardi , Désiré Sidibé , Samia Bouchafa

To enable large-scale reuse of real-world 3D assets, where garments and characters rarely share skeletons, templates, or dense correspondences, we present a fully automated virtual try-on system that dresses complex, multi-layer garments…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Cong Cao , Xianhang Cheng , Jingyuan Liu , Yujian Zheng , Zhenhui Lin , Ren Li , Meriem Chkir , Hao Li

This paper investigates a heterogeneous multi-vehicle, multi-modal sensing (H-MVMM) aided online precoding problem. The proposed H-MVMM scheme utilizes a vertical federated learning (VFL) framework to minimize pilot sequence length and…

Signal Processing · Electrical Eng. & Systems 2026-01-21 Haotian Zhang , Shijian Gao , Weibo Wen , Xiang Cheng , Liuqing Yang

Novel view synthesis (NVS) of in-the-wild garments is a challenging task due significant occlusions, complex human poses, and cloth deformations. Prior methods rely on synthetic 3D training data consisting of mostly unoccluded and static…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Johanna Karras , Yingwei Li , Yasamin Jafarian , Ira Kemelmacher-Shlizerman

Image-based virtual try-on aims to transfer target in-shop clothing to a dressed model image, the objectives of which are totally taking off original clothing while preserving the contents outside of the try-on area, naturally wearing…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Dan Song , Xuanpu Zhang , Jianhao Zeng , Pengxin Zhan , Qingguo Chen , Weihua Luo , An-An Liu

This paper introduces HapticVLM, a novel multimodal system that integrates vision-language reasoning with deep convolutional networks to enable real-time haptic feedback. HapticVLM leverages a ConvNeXt-based material recognition module to…

The Structure from Motion (SfM) challenge in computer vision is the process of recovering the 3D structure of a scene from a series of projective measurements that are calculated from a collection of 2D images, taken from different…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Joseph Rowell

Virtual clothes try-on has emerged as a vital feature in online shopping, offering consumers a critical tool to visualize how clothing fits. In our research, we introduce an innovative approach for virtual clothes try-on, utilizing a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Lingxiao Lu , Shengyi Wu , Haoxuan Sun , Junhong Gou , Jianlou Si , Chen Qian , Jianfu Zhang , Liqing Zhang
‹ Prev 1 8 9 10 Next ›