English
Related papers

Related papers: MMTryon: Multi-Modal Multi-Reference Control for H…

200 papers

Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot support low-latency and interactive garment control, which is crucial for applications such as…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Quanjian Song , Yefeng Shen , Mengting Chen , Hao Sun , Jinsong Lan , Xiaoyong Zhu , Bo Zheng , Liujuan Cao

Virtual try-on under arbitrary poses has attracted lots of research attention due to its huge potential applications. However, existing methods can hardly preserve the details in clothing texture and facial identity (face, hair) while…

Computer Vision and Pattern Recognition · Computer Science 2020-10-28 Jiahang Wang , Wei Zhang , Weizhong Liu , Tao Mei

In this paper, we introduce MRStyle, a comprehensive framework that enables color style transfer using multi-modality reference, including image and text. To achieve a unified style feature space for both modalities, we first develop a…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Jiancheng Huang , Yu Gao , Zequn Jie , Yujie Zhong , Xintong Han , Lin Ma

The development of smart cities has led to the generation of massive amounts of multi-modal data in the context of a range of tasks that enable a comprehensive monitoring of the smart city infrastructure and services. This paper surveys one…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Zhangyong Tang , Tianyang Xu , Xuefeng Zhu , Hui Li , Shaochuan Zhao , Tao Zhou , Chunyang Cheng , Xiaojun Wu , Josef Kittler

Image-based virtual try-on involves synthesizing perceptually convincing images of a model wearing a particular garment and has garnered significant research interest due to its immense practical applicability. Recent methods involve a two…

Computer Vision and Pattern Recognition · Computer Science 2021-09-16 Ayush Chopra , Rishabh Jain , Mayur Hemani , Balaji Krishnamurthy

Multimodal diffusion models for image editing generate outputs conditioned on both textual instructions and visual inputs, aiming to modify target regions while preserving the rest of the image. Although diffusion models have been shown to…

Cryptography and Security · Computer Science 2025-06-02 Ji Guo , Peihong Chen , Wenbo Jiang , Xiaolei Wen , Jiaming He , Jiachen Li , Guoming Lu , Aiguo Chen , Hongwei Li

Programming often involves converting detailed and complex specifications into code, a process during which developers typically utilize visual aids to more effectively convey concepts. While recent developments in Large Multimodal Models…

Computation and Language · Computer Science 2024-09-27 Kaixin Li , Yuchen Tian , Qisheng Hu , Ziyang Luo , Zhiyong Huang , Jing Ma

3D style transfer aims to generate stylized views of 3D scenes with specified styles, which requires high-quality generating and keeping multi-view consistency. Existing methods still suffer the challenges of high-quality stylization with…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Zijiang Yang , Zhongwei Qiu , Chang Xu , Dongmei Fu

Virtual try-off (VTOFF) aims to recover canonical flat-garment representations from images of dressed persons for standardized display and downstream virtual try-on. Prior methods often treat VTOFF as direct image translation driven by…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Shuang Liu , Ao Yu , Linkang Cheng , Xiwen Huang , Li Zhao , Junhui Liu , Zhiting Lin , Yu Liu

Virtual try-on aims to generate a photo-realistic fitting result given an in-shop garment and a reference person image. Existing methods usually build up multi-stage frameworks to deal with clothes warping and body blending respectively, or…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Shuai Bai , Huiling Zhou , Zhikang Li , Chang Zhou , Hongxia Yang

This paper introduces Multi-Garment Customized Model Generation, a unified framework based on Latent Diffusion Models (LDMs) aimed at addressing the unexplored task of synthesizing images with free combinations of multiple pieces of…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Yichen Liu , Penghui Du , Yi Liu Quanwei Zhang

The referring video object segmentation task (RVOS) involves segmentation of a text-referred object instance in the frames of a given video. Due to the complex nature of this multimodal task, which combines text reasoning, video…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Adam Botach , Evgenii Zheltonozhskii , Chaim Baskin

A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-modal tasks. In this paper, we introduce MMGen, a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Jiepeng Wang , Zhaoqing Wang , Hao Pan , Yuan Liu , Dongdong Yu , Changhu Wang , Wenping Wang

Image-based virtual try-on is one of the most promising applications of human-centric image generation due to its tremendous real-world potential. In this work, we take a step forwards to explore versatile virtual try-on solutions, which we…

Computer Vision and Pattern Recognition · Computer Science 2022-07-28 Zhenyu Xie , Zaiyu Huang , Fuwei Zhao , Haoye Dong , Michael Kampffmeyer , Xin Dong , Feida Zhu , Xiaodan Liang

Previous vision-language pre-training models mainly construct multi-modal inputs with tokens and objects (pixels) followed by performing cross-modality interaction between them. We argue that the input of only tokens and object features…

Computer Vision and Pattern Recognition · Computer Science 2022-09-15 Zejun Li , Zhihao Fan , Huaixiao Tou , Jingjing Chen , Zhongyu Wei , Xuanjing Huang

Fashion styling and personalized recommendations are pivotal in modern retail, contributing substantial economic value in the fashion industry. With the advent of vision-language models (VLM), new opportunities have emerged to enhance…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Kaicheng Pang , Xingxing Zou , Waikeung Wong

Video Visual Relation Detection (VidVRD) focuses on understanding how entities interact over time and space in videos, a key step for gaining deeper insights into video scenes beyond basic visual tasks. Traditional methods for VidVRD,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Xinjie Jiang , Chenxi Zheng , Xuemiao Xu , Bangzhen Liu , Weiying Zheng , Huaidong Zhang , Shengfeng He

This paper introduces a novel framework for virtual try-on, termed Wear-Any-Way. Different from previous methods, Wear-Any-Way is a customizable solution. Besides generating high-fidelity results, our method supports users to precisely…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Mengting Chen , Xi Chen , Zhonghua Zhai , Chen Ju , Xuewen Hong , Jinsong Lan , Shuai Xiao

Multi-object tracking (MOT) has profound applications in a variety of fields, including surveillance, sports analytics, self-driving, and cooperative robotics. Despite considerable advancements, existing MOT methodologies tend to falter…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Hamza Mukhtar , Muhammad Usman Ghani Khan

Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Question Answering (VQA),…

Artificial Intelligence · Computer Science 2024-03-04 Muhammad Arslan Manzoor , Sarah Albarri , Ziting Xian , Zaiqiao Meng , Preslav Nakov , Shangsong Liang
‹ Prev 1 8 9 10 Next ›