English
Related papers

Related papers: ConsistentID: Portrait Generation with Multimodal …

200 papers

Person re-identification (re-ID) concerns the matching of subject images across different camera views in a multi camera surveillance system. One of the major challenges in person re-ID is pose variations across the camera network, which…

Computer Vision and Pattern Recognition · Computer Science 2021-04-29 Amena Khatun , Simon Denman , Sridha Sridharan , Clinton Fookes

Diffusion Purification, purifying noised images with diffusion models, has been widely used for enhancing certified robustness via randomized smoothing. However, existing frameworks often grapple with the balance between efficiency and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Yiquan Li , Zhongzhu Chen , Kun Jin , Jiongxiao Wang , Bo Li , Chaowei Xiao

Multi-person identity-preserving generation requires binding multiple reference faces to specified locations under a text prompt. Strong identity/layout conditions often trigger copy-paste shortcuts and weaken prompt-driven controllability.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Longhui Yuan

We investigate how to generate multimodal image outputs, such as RGB, depth, and surface normals, with a single generative model. The challenge is to produce outputs that are realistic, and also consistent with each other. Our solution…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Zhen Zhu , Yijun Li , Weijie Lyu , Krishna Kumar Singh , Zhixin Shu , Soeren Pirk , Derek Hoiem

Recently, zero-shot methods like InstantID have revolutionized identity-preserving generation. Unlike multi-image finetuning approaches such as DreamBooth, these zero-shot methods leverage powerful facial encoders to extract identity…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Yiren Song , Pei Yang , Hai Ci , Mike Zheng Shou

We address the problem of learning person-specific facial priors from a small number (e.g., 20) of portrait photos of the same person. This enables us to edit this specific person's facial appearance, such as expression and lighting, while…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Zheng Ding , Xuaner Zhang , Zhihao Xia , Lars Jebe , Zhuowen Tu , Xiuming Zhang

Recent diffusion-based Single-image 3D portrait generation methods typically employ 2D diffusion models to provide multi-view knowledge, which is then distilled into 3D representations. However, these methods usually struggle to produce…

Computer Vision and Pattern Recognition · Computer Science 2024-11-18 Haoran Wei , Wencheng Han , Xingping Dong , Jianbing Shen

Multi-view generation with camera pose control and prompt-based customization are both essential elements for achieving controllable generative models. However, existing multi-view generation models do not support customization with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Minjung Shin , Hyunin Cho , Sooyeon Go , Jin-Hwa Kim , Youngjung Uh

Nowadays, deep learning models have reached incredible performance in the task of image generation. Plenty of literature works address the task of face generation and editing, with human and automatic systems that struggle to distinguish…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Giuseppe Tarollo , Tomaso Fontanini , Claudio Ferrari , Guido Borghi , Andrea Prati

The emergence of generative models enables the creation of texts and images tailored to users' preferences. Existing personalized generative models have two critical limitations: lacking a dedicated paradigm for accurate preference…

Information Retrieval · Computer Science 2026-04-23 Yuting Zhang , Ying Sun , Dazhong Shen , Ziwei Xie , Feng Liu , Changwang Zhang , Xiang Liu , Jun Wang , Hui Xiong

Current diffusion models for human image animation struggle to ensure identity (ID) consistency. This paper presents StableAnimator, the first end-to-end ID-preserving video diffusion framework, which synthesizes high-quality videos without…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Shuyuan Tu , Zhen Xing , Xintong Han , Zhi-Qi Cheng , Qi Dai , Chong Luo , Zuxuan Wu

We present a deep learning-based framework for portrait reenactment from a single picture of a target (one-shot) and a video of a driving subject. Existing facial reenactment methods suffer from identity mismatch and produce inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2020-04-28 Sitao Xiang , Yuming Gu , Pengda Xiang , Mingming He , Koki Nagano , Haiwei Chen , Hao Li

Video identity customization seeks to produce high-fidelity videos that maintain consistent identity and exhibit significant dynamics based on users' reference images. However, existing approaches face two key challenges: identity…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Hengjia Li , Lifan Jiang , Xi Xiao , Tianyang Wang , Hongwei Yi , Boxi Wu , Deng Cai

In human-centric content generation, the pre-trained text-to-image models struggle to produce user-wanted portrait images, which retain the identity of individuals while exhibiting diverse expressions. This paper introduces our efforts…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Renshuai Liu , Bowen Ma , Wei Zhang , Zhipeng Hu , Changjie Fan , Tangjie Lv , Yu Ding , Xuan Cheng

Facial stylization aims to transform facial images into appealing, high-quality stylized portraits, with the critical challenge of accurately learning the target style while maintaining content consistency with the original image. Although…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Zhanyi Lu , Yue Zhou

Due to the data-driven nature of current face identity (FaceID) customization methods, all state-of-the-art models rely on large-scale datasets containing millions of high-quality text-image pairs for training. However, none of these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Shuhe Wang , Xiaoya Li , Jiwei Li , Guoyin Wang , Xiaofei Sun , Bob Zhu , Han Qiu , Mo Yu , Shengjie Shen , Tianwei Zhang , Eduard Hovy

We present Follow-Your-Emoji-Faster, an efficient diffusion-based framework for freestyle portrait animation driven by facial landmarks. The main challenges in this task are preserving the identity of the reference portrait, accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Yue Ma , Zexuan Yan , Hongyu Liu , Hongfa Wang , Heng Pan , Yingqing He , Junkun Yuan , Ailing Zeng , Chengfei Cai , Heung-Yeung Shum , Zhifeng Li , Wei Liu , Linfeng Zhang , Qifeng Chen

Recent advances in diffusion models have significantly improved text-to-face generation, but achieving fine-grained control over facial features remains a challenge. Existing methods often require training additional modules to handle…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Liang Shi , Yun Fu

State-of-the-art deep neural network models have reached near perfect face recognition accuracy rates on controlled high-resolution face images. However, their performance is drastically degraded when they are tested with very…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Vahid Reza Khazaie , Nicky Bayat , Yalda Mohsenzadeh

Face aging or de-aging with generative AI has gained significant attention for its applications in such fields like forensics, security, and media. However, most state of the art methods rely on conditional Generative Adversarial Networks…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Luis S. Luevano , Pavel Korshunov , Sebastien Marcel