中文
相关论文

相关论文: CosmicMan: A Text-to-Image Foundation Model for Hu…

200 篇论文

Customized text-to-image generation, which synthesizes images based on user-specified concepts, has made significant progress in handling individual concepts. However, when extended to multiple concepts, existing methods often struggle with…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Jiaxiu Jiang , Yabo Zhang , Kailai Feng , Xiaohe Wu , Wenbo Li , Renjing Pei , Fan Li , Wangmeng Zuo

Recent advancements in controllable human image generation have led to zero-shot generation using structural signals (e.g., pose, depth) or facial appearance. Yet, generating human images conditioned on multiple parts of human appearance…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Zehuan Huang , Hongxing Fan , Lipeng Wang , Lu Sheng

Human image editing includes tasks like changing a person's pose, their clothing, or editing the image according to a text prompt. However, prior work often tackles these tasks separately, overlooking the benefit of mutual reinforcement…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Nannan Li , Qing Liu , Krishna Kumar Singh , Yilin Wang , Jianming Zhang , Bryan A. Plummer , Zhe Lin

Estimation of human shape and pose from a single image is a challenging task. It is an even more difficult problem to map the identified human shape onto a 3D human model. Existing methods map manually labelled human pixels in real 2D…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Mithun Lal , Anthony Paproki , Nariman Habili , Lars Petersson , Olivier Salvado , Clinton Fookes

In recent years, general visual foundation models (VFMs) have witnessed increasing adoption, particularly as image encoders for popular multi-modal large language models (MLLMs). However, without semantically fine-grained supervision, these…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Tongkun Guan , Zining Wang , Pei Fu , Zhengtao Guo , Wei Shen , Kai Zhou , Tiezhu Yue , Chen Duan , Hao Sun , Qianyi Jiang , Junfeng Luo , Xiaokang Yang

Image descriptions can help visually impaired people to quickly understand the image content. While we made significant progress in automatically describing images and optical character recognition, current approaches are unable to include…

计算机视觉与模式识别 · 计算机科学 2020-08-05 Oleksii Sidorov , Ronghang Hu , Marcus Rohrbach , Amanpreet Singh

Image-text matching has been a long-standing problem, which seeks to connect vision and language through semantic understanding. Due to the capability to manage large-scale raw data, unsupervised hashing-based approaches have gained…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Fan Zhang , Xian-Sheng Hua , Chong Chen , Xiao Luo

We consider the problem of obese human mesh recovery, i.e., fitting a parametric human mesh to images of obese people. Despite obese person mesh fitting being an important problem with numerous applications (e.g., healthcare), much recent…

计算机视觉与模式识别 · 计算机科学 2021-07-14 Ren Li , Meng Zheng , Srikrishna Karanam , Terrence Chen , Ziyan Wu

Text-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text inputs. This paper introduces StickMotion, an efficient…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Tao Wang , Zhihua Wu , Qiaozhi He , Jiaming Chu , Ling Qian , Yu Cheng , Junliang Xing , Jian Zhao , Lei Jin

While large text-to-image models are able to synthesize "novel" images, these images are necessarily a reflection of the training data. The problem of data attribution in such models -- which of the images in the training set are most…

计算机视觉与模式识别 · 计算机科学 2023-08-09 Sheng-Yu Wang , Alexei A. Efros , Jun-Yan Zhu , Richard Zhang

Revolutionary advancements in text-to-image models have unlocked new dimensions for sophisticated content creation, such as text-conditioned image editing, enabling the modification of existing images based on textual guidance. This…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Haoyu Zheng , Wenqiao Zhang , Yaoke Wang , Juncheng Li , Zheqi Lv , Xin Min , Mengze Li , Dongping Zhang , Siliang Tang , Yueting Zhuang

Existing text-to-image synthesis methods generally are only applicable to words in the training dataset. However, human faces are so variable to be described with limited words. So this paper proposes the first free-style text-to-face…

计算机视觉与模式识别 · 计算机科学 2022-03-30 Jianxin Sun , Qiyao Deng , Qi Li , Muyi Sun , Min Ren , Zhenan Sun

Image generation has achieved remarkable progress with the development of large-scale text-to-image models, especially diffusion-based models. However, generating human images with plausible details, such as faces or hands, remains…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Yuxuan Wang , Tianwei Cao , Huayu Zhang , Zhongjiang He , Kongming Liang , Zhanyu Ma

Generating dense multiview images from text prompts is crucial for creating high-fidelity 3D assets. Nevertheless, existing methods struggle with space-view correspondences, resulting in sparse and low-quality outputs. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Bonan Li , Zicheng Zhang , Xingyi Yang , Xinchao Wang

Recent advancements in text-to-image generation using diffusion models have significantly improved the quality of generated images and expanded the ability to depict a wide range of objects. However, ensuring that these models adhere…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Michail Tarasiou , Stylianos Moschoglou , Jiankang Deng , Stefanos Zafeiriou

Despite recent advancements in text-to-image models, achieving semantically accurate images in text-to-image diffusion models is a persistent challenge. While existing initial latent optimization methods have demonstrated impressive…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Aravindan Sundaram , Ujjayan Pal , Abhimanyu Chauhan , Aishwarya Agarwal , Srikrishna Karanam

While foundation models have achieved remarkable results across a diversity of domains, they still rely on human-generated data, such as text, as a fundamental source of knowledge. However, this data is ultimately the product of human…

神经元与认知 · 定量生物学 2026-01-21 Maël Donoso

Generating an image from a provided descriptive text is quite a challenging task because of the difficulty in incorporating perceptual information (object shapes, colors, and their interactions) along with providing high relevancy related…

计算机视觉与模式识别 · 计算机科学 2020-07-03 Kanish Garg , Ajeet kumar Singh , Dorien Herremans , Brejesh Lall

The customization of text-to-image models has seen significant advancements, yet generating multiple personalized concepts remains a challenging task. Current methods struggle with attribute leakage and layout confusion when handling…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Zebin Yao , Fangxiang Feng , Ruifan Li , Xiaojie Wang

Recent vision-language foundation models still frequently produce outputs misaligned with their inputs, evidenced by object hallucination in captioning and prompt misalignment in the text-to-image generation model. Recent studies have…

计算机视觉与模式识别 · 计算机科学 2024-12-25 JeongYeon Nam , Jinbae Im , Wonjae Kim , Taeho Kil
‹ 上一页 1 8 9 10 下一页 ›