English
Related papers

Related papers: PuLID: Pure and Lightning ID Customization via Con…

200 papers

Text-to-image generation for personalized identities aims at incorporating the specific identity into images using a text prompt and an identity image. Based on the powerful generative capabilities of DDPMs, many previous works adopt…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Jinyu Gu , Haipeng Liu , Meng Wang , Yang Wang

While adversarial attacks on vision-and-language pretraining (VLP) models have been explored, generating natural adversarial samples crafted through realistic and semantically meaningful perturbations remains an open challenge. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Ying Yang , Jie Zhang , Xiao Lv , Di Lin , Tao Xiang , Qing Guo

The remarkable advancement in text-to-image generation models significantly boosts the research in ID customization generation. However, existing personalization methods cannot simultaneously satisfy high fidelity and high-efficiency…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Peng Xing , Ning Wang , Jianbo Ouyang , Zechao Li

Personalized diffusion models have shown remarkable success in Text-to-Image (T2I) generation by enabling the injection of user-defined concepts into diverse contexts. However, balancing concept fidelity with contextual alignment remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Shamil Ayupov , Maksim Nakhodnov , Anastasia Yaschenko , Andrey Kuznetsov , Aibek Alanov

Digital identity verification systems used in remote onboarding rely on document images to authenticate users, making them vulnerable to localized manipulations of key identity fields such as facial photographs and textual information.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Abhishek Kumar , Riya Tapwal , Carsten Maple , Mark Hooper

While Diffusion Large Language Models (DLLMs) have demonstrated remarkable capabilities in multi-modal generation, performing precise, training-free image editing remains an open challenge. Unlike continuous diffusion models, the discrete…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Zifeng Zhu , Jiaming Han , Jiaxiang Zhao , Minnan Luo , Xiangyu Yue

We present LiDAR-EDIT, a novel paradigm for generating synthetic LiDAR data for autonomous driving. Our framework edits real-world LiDAR scans by introducing new object layouts while preserving the realism of the background environment.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Shing-Hei Ho , Bao Thach , Minghan Zhu

Tuning-free approaches adapting large-scale pre-trained video diffusion models for identity-preserving text-to-video generation (IPT2V) have gained popularity recently due to their efficacy and scalability. However, significant challenges…

Graphics · Computer Science 2025-02-26 Yunpeng Zhang , Qiang Wang , Fan Jiang , Yaqi Fan , Mu Xu , Yonggang Qi

Image personalization has garnered attention for its ability to customize Text-to-Image generation using only a few reference images. However, a key challenge in image personalization is the issue of conceptual coupling, where the limited…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Shizhan Liu , Hao Zheng , Hang Yu , Jianguo Li

The few-shot fine-tuning of Latent Diffusion Models (LDMs) has enabled them to grasp new concepts from a limited number of images. However, given the vast amount of personal images accessible online, this capability raises critical concerns…

Cryptography and Security · Computer Science 2024-06-24 Ang Li , Yichuan Mo , Mingjie Li , Yisen Wang

Image inpainting aims at restoring missing regions of corrupted images, which has many applications such as image restoration and object removal. However, current GAN-based generative inpainting models do not explicitly exploit the…

Computer Vision and Pattern Recognition · Computer Science 2020-09-15 Ang Li , Jianzhong Qi , Rui Zhang , Xingjun Ma , Kotagiri Ramamohanarao

Text-to-image diffusion models have recently received increasing interest for their astonishing ability to produce high-fidelity images from solely text inputs. Subsequent research efforts aim to exploit and apply their capabilities to real…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Manuel Brack , Felix Friedrich , Katharina Kornmeier , Linoy Tsaban , Patrick Schramowski , Kristian Kersting , Apolinário Passos

Text-to-image (T2I) diffusion models have become prominent tools for generating high-fidelity images from text prompts. However, when trained on unfiltered internet data, these models can produce unsafe, incorrect, or stylistically…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Rohit Jena , Ali Taghibakhshi , Sahil Jain , Gerald Shen , Nima Tajbakhsh , Arash Vahdat

Recent advancements in adapting vision-language pre-training models like CLIP for person re-identification (ReID) tasks often rely on complex adapter design or modality-specific tuning while neglecting cross-modal interaction, leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Yunfei Xie , Yuxuan Cheng , Juncheng Wu , Haoyu Zhang , Yuyin Zhou , Shoudong Han

To reduce the latency associated with autoretrogressive LLM inference, speculative decoding has emerged as a novel decoding paradigm, where future tokens are drafted and verified in parallel. However, the practical deployment of speculative…

Computation and Language · Computer Science 2024-12-03 Shwetha Somasundaram , Anirudh Phukan , Apoorv Saxena

In recent years, the fusion of camera data with LiDAR measurements has emerged as a powerful approach to enhance spatial understanding. This study introduces a novel, hardware-agnostic methodology that generates colourised point clouds from…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Pasindu Ranasinghe , Dibyayan Patra , Bikram Banerjee , Simit Raval

Generative recommendation systems have achieved significant advances by leveraging semantic IDs to represent items. However, existing approaches that tokenize each modality independently face two critical limitations: (1) redundancy across…

Information Retrieval · Computer Science 2026-01-14 Haven Kim , Yupeng Hou , Julian McAuley

Diffusion-based technologies have made significant strides, particularly in personalized and customized facialgeneration. However, existing methods face challenges in achieving high-fidelity and detailed identity (ID)consistency, primarily…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Jiehui Huang , Xiao Dong , Wenhui Song , Zheng Chong , Zhenchao Tang , Jun Zhou , Yuhao Cheng , Long Chen , Hanhui Li , Yiqiang Yan , Shengcai Liao , Xiaodan Liang

Benefited from image-text contrastive learning, pre-trained vision-language models, e.g., CLIP, allow to direct leverage texts as images (TaI) for parameter-efficient fine-tuning (PEFT). While CLIP is capable of making image features to be…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Chun-Mei Feng , Kai Yu , Xinxing Xu , Salman Khan , Rick Siow Mong Goh , Wangmeng Zuo , Yong Liu

Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images according to prompt alignment and perceptual quality. Existing reward models are commonly…

Artificial Intelligence · Computer Science 2026-05-22 Kuei-Chun Kao , Daixuan Huo , Yuanhao Ban , Cho-Jui Hsieh