中文
相关论文

相关论文: UniPortrait: A Unified Framework for Identity-Pres…

200 篇论文

We concentrate on a novel human-centric image synthesis task, that is, given only one reference facial photograph, it is expected to generate specific individual images with diverse head positions, poses, facial expressions, and…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Chao Liang , Fan Ma , Linchao Zhu , Yingying Deng , Yi Yang

Human insertion aims to naturally place specific individuals into a target background. Although existing image editing models may have such ability, they often produce failure cases, including inappropriate human pose in new background,…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Jie Li , Shulian Zhang , Yangyang Gao , Wenbo Li , Yulun Zhang , Yong Guo , Jian Chen

Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when referencing details across multiple input images. In this work,…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Pengcheng Xu , Peng Tang , Donghao Luo , Xiaobin Hu , Weichu Cui , Qingdong He , Zhennan Chen , Jiangning Zhang , Charles Ling , Boyu Wang

Facial personalization faces challenges to maintain identity fidelity without disrupting the foundation model's prompt consistency. The mainstream personalization models employ identity embedding to integrate identity information within the…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Yiyang Cai , Zhengkai Jiang , Yulong Liu , Chunyang Jiang , Wei Xue , Yike Guo , Wenhan Luo

The latest developments in Face Restoration have yielded significant advancements in visual quality through the utilization of diverse diffusion priors. Nevertheless, the uncertainty of face identity introduced by identity-obscure inputs…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Yushun Fang , Lu Liu , Xiang Gao , Qiang Hu , Ning Cao , Jianghe Cui , Gang Chen , Xiaoyun Zhang

Achieving flexible and high-fidelity identity-preserved image generation remains formidable, particularly with advanced Diffusion Transformers (DiTs) like FLUX. We introduce InfiniteYou (InfU), one of the earliest robust frameworks…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Liming Jiang , Qing Yan , Yumin Jia , Zichuan Liu , Hao Kang , Xin Lu

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: the model, the tasks,…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Chi Zhang , Jiepeng Wang , Youming Wang , Yuanzhi Liang , Xiaoyan Yang , Zuoxin Li , Haibin Huang , Xuelong Li

Face reenactment and portrait relighting are essential tasks in portrait editing, yet they are typically addressed independently, without much synergy. Most face reenactment methods prioritize motion control and multiview consistency, while…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Yizhou Zhao , Chunjiang Liu , Haoyu Chen , Bhiksha Raj , Min Xu , Tadas Baltrusaitis , Mitch Rundle , HsiangTao Wu , Kamran Ghasedi

This work presents FaceX framework, a novel facial generalist model capable of handling diverse facial tasks simultaneously. To achieve this goal, we initially formulate a unified facial representation for a broad spectrum of facial editing…

计算机视觉与模式识别 · 计算机科学 2024-01-02 Yue Han , Jiangning Zhang , Junwei Zhu , Xiangtai Li , Yanhao Ge , Wei Li , Chengjie Wang , Yong Liu , Xiaoming Liu , Ying Tai

Daily monitoring of intra-personal facial changes associated with health and emotional conditions has great potential to be useful for medical, healthcare, and emotion recognition fields. However, the approach for capturing intra-personal…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Yusuke Akamatsu , Terumi Umematsu , Hitoshi Imaoka , Shizuko Gomi , Hideo Tsurushima

Recent advances in text-to-image generation have made remarkable progress in synthesizing realistic human photos conditioned on given text prompts. However, existing personalized generation methods cannot simultaneously satisfy the…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Zhen Li , Mingdeng Cao , Xintao Wang , Zhongang Qi , Ming-Ming Cheng , Ying Shan

Recently, diffusion models have made significant strides in synthesizing realistic 2D human images based on provided text prompts. Building upon this, researchers have extended 2D text-to-image diffusion models into the 3D domain for…

计算机视觉与模式识别 · 计算机科学 2024-08-12 Weijie Wang , Jichao Zhang , Chang Liu , Xia Li , Xingqian Xu , Humphrey Shi , Nicu Sebe , Bruno Lepri

Recent advances in generative modeling have enabled the generation of high-quality synthetic data that is applicable in a variety of domains, including face recognition. Here, state-of-the-art generative models typically rely on…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Darian Tomašević , Fadi Boutros , Chenhao Lin , Naser Damer , Vitomir Štruc , Peter Peer

Large-scale text-to-image models including Stable Diffusion are capable of generating high-fidelity photorealistic portrait images. There is an active research area dedicated to personalizing these models, aiming to synthesize specific…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Junha Hyung , Jaeyo Shin , Jaegul Choo

Recent advancements in controllable human image generation have led to zero-shot generation using structural signals (e.g., pose, depth) or facial appearance. Yet, generating human images conditioned on multiple parts of human appearance…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Zehuan Huang , Hongxing Fan , Lipeng Wang , Lu Sheng

Group portrait editing is highly desirable since users constantly want to add a person, delete a person, or manipulate existing persons. It is also challenging due to the intricate dynamics of human interactions and the diverse gestures. In…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Yuming Jiang , Nanxuan Zhao , Qing Liu , Krishna Kumar Singh , Shuai Yang , Chen Change Loy , Ziwei Liu

Human face generation and editing represent an essential task in the era of computer vision and the digital world. Recent studies have shown remarkable progress in multi-modal face generation and editing, for instance, using face…

计算机视觉与模式识别 · 计算机科学 2024-02-06 Mohammadreza Mofayezi , Reza Alipour , Mohammad Ali Kakavand , Ehsaneddin Asgari

The fashion domain encompasses a variety of real-world multimodal tasks, including multimodal retrieval and multimodal generation. The rapid advancements in artificial intelligence generated content, particularly in technologies like large…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Xiangyu Zhao , Yuehan Zhang , Wenlong Zhang , Xiao-Ming Wu

Current diffusion-based acceleration methods for long-portrait animation struggle to ensure identity (ID) consistency. This paper presents FlashPortrait, an end-to-end video diffusion transformer capable of synthesizing ID-preserving,…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Shuyuan Tu , Yueming Pan , Yinming Huang , Xintong Han , Zhen Xing , Qi Dai , Kai Qiu , Chong Luo , Zuxuan Wu

Portrait Animation aims to synthesize a lifelike video from a single source image, using it as an appearance reference, with motion (i.e., facial expressions and head pose) derived from a driving video, audio, text, or generation. Instead…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Jianzhu Guo , Dingyun Zhang , Xiaoqiang Liu , Zhizhou Zhong , Yuan Zhang , Pengfei Wan , Di Zhang