English
Related papers

Related papers: Visual Persona: Foundation Model for Full-Body Hum…

200 papers

Text-to-image personalization aims to teach a pre-trained diffusion model to reason about novel, user provided concepts, embedding them into new scenes guided by natural language prompts. However, current personalization approaches struggle…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Rinon Gal , Moab Arar , Yuval Atzmon , Amit H. Bermano , Gal Chechik , Daniel Cohen-Or

Human re-rendering from a single image is a starkly under-constrained problem, and state-of-the-art algorithms often exhibit undesired artefacts, such as over-smoothing, unrealistic distortions of the body parts and garments, or implausible…

Computer Vision and Pattern Recognition · Computer Science 2021-01-12 Kripasindhu Sarkar , Dushyant Mehta , Weipeng Xu , Vladislav Golyanik , Christian Theobalt

To make 3D human avatars widely available, we must be able to generate a variety of 3D virtual humans with varied identities and shapes in arbitrary poses. This task is challenging due to the diversity of clothed body shapes, their complex…

Computer Vision and Pattern Recognition · Computer Science 2022-04-14 Xu Chen , Tianjian Jiang , Jie Song , Jinlong Yang , Michael J. Black , Andreas Geiger , Otmar Hilliges

Recent advancements in controllable human image generation have led to zero-shot generation using structural signals (e.g., pose, depth) or facial appearance. Yet, generating human images conditioned on multiple parts of human appearance…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Zehuan Huang , Hongxing Fan , Lipeng Wang , Lu Sheng

Visual storytelling is the task of generating stories based on a sequence of images. Inspired by the recent works in neural generation focusing on controlling the form of text, this paper explores the idea of generating these stories in…

Computation and Language · Computer Science 2019-06-18 Shrimai Prabhumoye , Khyathi Raghavi Chandu , Ruslan Salakhutdinov , Alan W Black

While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centric images, an intractable problem is how to preserve the face identity for conditioned face images. Existing methods either require…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Zhuowei Chen , Shancheng Fang , Wei Liu , Qian He , Mengqi Huang , Yongdong Zhang , Zhendong Mao

Recent advances in text-to-image generation have enabled the creation of high-quality images with diverse applications. However, accurately describing desired visual attributes can be challenging, especially for non-experts in art and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Tong Wu , Yinghao Xu , Ryan Po , Mengchen Zhang , Guandao Yang , Jiaqi Wang , Ziwei Liu , Dahua Lin , Gordon Wetzstein

Video personalization aims to generate videos that faithfully reflect a user-provided subject while following a text prompt. However, existing approaches often rely on heavy video-based finetuning or large-scale video datasets, which impose…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Hyunkoo Lee , Wooseok Jang , Jini Yang , Taehwan Kim , Sangoh Kim , Sangwon Jung , Seungryong Kim

This paper presents a novel method to manipulate the visual appearance (pose and attribute) of a person image according to natural language descriptions. Our method can be boiled down to two stages: 1) text guided pose generation and 2)…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Xingran Zhou , Siyu Huang , Bin Li , Yingming Li , Jiachen Li , Zhongfei Zhang

Driving a high-quality and photorealistic full-body virtual human from a few RGB cameras is a challenging problem that has become increasingly relevant with emerging virtual reality technologies. A promising solution to democratize such…

Image and Video Processing · Electrical Eng. & Systems 2025-08-26 Anton Zubekhin , Heming Zhu , Paulo Gotardo , Thabo Beeler , Marc Habermann , Christian Theobalt

Leveraging Stable Diffusion for the generation of personalized portraits has emerged as a powerful and noteworthy tool, enabling users to create high-fidelity, custom character avatars based on their specific prompts. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Siying Cui , Jia Guo , Xiang An , Jiankang Deng , Yongle Zhao , Xinyu Wei , Ziyong Feng

Massive captured face images are stored in the database for the identification of individuals. However, these images can be observed unintentionally by data managers, which is not at the will of individuals and may cause privacy violations.…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Tao Wang , Yushu Zhang , Zixuan Yang , Xiangli Xiao , Hua Zhang , Zhongyun Hua

In this paper, we present an end-to-end approach to generate high-resolution person images conditioned on texts only. State-of-the-art text-to-image generation models are mainly designed for center-object generation, e.g., flowers and…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Deyin Liu , Lin Yuanbo Wu , Bo Li , Zongyuan Ge

The goal of image-based virtual try-on is to generate an image of the target person naturally wearing the given clothing. However, existing methods solely focus on the frontal try-on using the frontal clothing. When the views of the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Haoyu Wang , Zhilu Zhang , Donglin Di , Shiliang Zhang , Wangmeng Zuo

Vanilla text-to-image diffusion models struggle with generating accurate human images, commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limbs.Existing methods address this issue mostly by fine-tuning…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Junyan Wang , Zhenhong Sun , Zhiyu Tan , Xuanbai Chen , Weihua Chen , Hao Li , Cheng Zhang , Yang Song

Recently, personalized portrait generation with a text-to-image diffusion model has significantly advanced with Textual Inversion, emerging as a promising approach for creating high-fidelity personalized images. Despite its potential,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Hyun-Jun Jin , Young-Eun Kim , Seong-Whan Lee

Human image editing includes tasks like changing a person's pose, their clothing, or editing the image according to a text prompt. However, prior work often tackles these tasks separately, overlooking the benefit of mutual reinforcement…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Nannan Li , Qing Liu , Krishna Kumar Singh , Yilin Wang , Jianming Zhang , Bryan A. Plummer , Zhe Lin

The goal of human stylization is to transfer full-body human photos to a style specified by a single art character reference image. Although previous work has succeeded in example-based stylization of faces and generic scenes, full-body…

Computer Vision and Pattern Recognition · Computer Science 2023-04-17 Aiyu Cui , Svetlana Lazebnik

Visual concept personalization aims to transfer only specific image attributes, such as identity, expression, lighting, and style, into unseen contexts. However, existing methods rely on holistic embeddings from general-purpose image…

Person re-identification (re-ID) is a challenging problem especially when no labels are available for training. Although recent deep re-ID methods have achieved great improvement, it is still difficult to optimize deep re-ID model without…

Computer Vision and Pattern Recognition · Computer Science 2018-11-07 Fengxiang Yang , Zhun Zhong , Zhiming Luo , Sheng Lian , Shaozi Li