English
Related papers

Related papers: CosmicMan: A Text-to-Image Foundation Model for Hu…

200 papers

Building on the success of diffusion models, significant advancements have been made in multimodal image generation tasks. Among these, human image generation has emerged as a promising technique, offering the potential to revolutionize the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Shiyue Zhang , Zheng Chong , Xi Lu , Wenqing Zhang , Haoxiang Li , Xujie Zhang , Jiehui Huang , Xiao Dong , Xiaodan Liang

Recent text-to-image generation methods provide a simple yet exciting conversion capability between text and image domains. While these methods have incrementally improved the generated image fidelity and text relevancy, several pivotal…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Oran Gafni , Adam Polyak , Oron Ashual , Shelly Sheynin , Devi Parikh , Yaniv Taigman

Text-to-image synthesis has achieved high-quality results with recent advances in diffusion models. However, text input alone has high spatial ambiguity and limited user controllability. Most existing methods allow spatial control through…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Yuki Endo

We propose a data-driven approach for context-aware person image generation. Specifically, we attempt to generate a person image such that the synthesized instance can blend into a complex scene. In our method, the position, scale, and…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Prasun Roy , Saumik Bhattacharya , Subhankar Ghosh , Umapada Pal , Michael Blumenstein

We propose a zero-shot approach for generating consistent videos of animated characters based on Text-to-Image (T2I) diffusion models. Existing Text-to-Video (T2V) methods are expensive to train and require large-scale video datasets to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Abdelrahman Eldesokey , Peter Wonka

Human generation has achieved significant progress. Nonetheless, existing methods still struggle to synthesize specific regions such as faces and hands. We argue that the main reason is rooted in the training data. A holistic human dataset…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Jianglin Fu , Shikai Li , Yuming Jiang , Kwan-Yee Lin , Wayne Wu , Ziwei Liu

Despite significant advances in large-scale text-to-image models, achieving hyper-realistic human image generation remains a desirable yet unsolved task. Existing models like Stable Diffusion and DALL-E 2 tend to generate human images with…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Xian Liu , Jian Ren , Aliaksandr Siarohin , Ivan Skorokhodov , Yanyu Li , Dahua Lin , Xihui Liu , Ziwei Liu , Sergey Tulyakov

We present BootComp, a novel framework based on text-to-image diffusion models for controllable human image generation with multiple reference garments. Here, the main bottleneck is data acquisition for training: collecting a large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Yisol Choi , Sangkyung Kwak , Sihyun Yu , Hyungwon Choi , Jinwoo Shin

Anonymization plays a key role in protecting sensible information of individuals in real world datasets. Self-driving cars for example need high resolution facial features to track people and their viewing direction to predict future…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Pascal Zwick , Kevin Roesch , Marvin Klemp , Oliver Bringmann

Text-driven large scene image synthesis has made significant progress with diffusion models, but controlling it is challenging. While using additional spatial controls with corresponding texts has improved the controllability of large scene…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Gwanghyun Kim , Dong Un Kang , Hoigi Seo , Hayeon Kim , Se Young Chun

Text-to-image diffusion models have significantly advanced in conditional image generation. However, these models usually struggle with accurately rendering images featuring humans, resulting in distorted limbs and other anomalies. This…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Guian Fang , Wenbiao Yan , Yuanfan Guo , Jianhua Han , Zutao Jiang , Hang Xu , Shengcai Liao , Xiaodan Liang

Existing text-to-image diffusion models struggle to synthesize realistic images given dense captions, where each text prompt provides a detailed description for a specific image region. To address this, we propose DenseDiffusion, a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Yunji Kim , Jiyoung Lee , Jin-Hwa Kim , Jung-Woo Ha , Jun-Yan Zhu

Training data is at the core of any successful text-to-image models. The quality and descriptiveness of image text are crucial to a model's performance. Given the noisiness and inconsistency in web-scraped datasets, recent works shifted…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Manuel Brack , Sudeep Katakol , Felix Friedrich , Patrick Schramowski , Hareesh Ravi , Kristian Kersting , Ajinkya Kale

People get informed of a daily task plan through diverse media involving both texts and images. However, most prior research only focuses on LLM's capability of textual plan generation. The potential of large-scale models in providing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Xiaoxin Lu , Ranran Haoran Zhang , Yusen Zhang , Rui Zhang

Large text-to-image models have revolutionized the ability to generate imagery using natural language. However, particularly unique or personal visual concepts, such as pets and furniture, will not be captured by the original model. This…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Xingzhe He , Zhiwen Cao , Nicholas Kolkin , Lantao Yu , Kun Wan , Helge Rhodin , Ratheesh Kalarot

This paper focuses on creating synthetic data to improve the quality of image captions. Existing works typically have two shortcomings. First, they caption images from scratch, ignoring existing alt-text metadata, and second, lack…

Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likeness to capture a character's authentic essence. Their motions typically synchronize with low-level cues like audio rhythm,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Jianwen Jiang , Weihong Zeng , Zerong Zheng , Jiaqi Yang , Chao Liang , Wang Liao , Han Liang , Yuan Zhang , Mingyuan Gao

We present a novel framework, InfinityGAN, for arbitrary-sized image generation. The task is associated with several key challenges. First, scaling existing models to an arbitrarily large image size is resource-constrained, in terms of both…

Computer Vision and Pattern Recognition · Computer Science 2022-03-14 Chieh Hubert Lin , Hsin-Ying Lee , Yen-Chi Cheng , Sergey Tulyakov , Ming-Hsuan Yang

Text-to-image generation models have achieved remarkable advancements in recent years, aiming to produce realistic images from textual descriptions. However, these models often struggle with generating anatomically accurate representations…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Haozhuo Zhang , Bin Zhu , Yu Cao , Yanbin Hao

Generating high-fidelity images of humans with fine-grained control over attributes such as hairstyle and clothing remains a core challenge in personalized text-to-image synthesis. While prior methods emphasize identity preservation from a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Guocheng Gordon Qian , Daniil Ostashev , Egor Nemchinov , Avihay Assouline , Sergey Tulyakov , Kuan-Chieh Jackson Wang , Kfir Aberman