中文
相关论文

相关论文: Consistent text-to-image generation via scene de-c…

200 篇论文

Text-to-LiDAR generation can customize 3D data with rich structures and diverse scenes for downstream tasks. However, the scarcity of Text-LiDAR pairs often causes insufficient training priors, generating overly smooth 3D scenes. Moreover,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Wentao Qu , Guofeng Mei , Yang Wu , Yongshun Gong , Xiaoshui Huang , Liang Xiao

Text-to-image (T2I) diffusion models, with their impressive generative capabilities, have been adopted for image editing tasks, demonstrating remarkable efficacy. However, due to attention leakage and collision between the cross-attention…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Xingxi Yin , Zhi Li , Jingfeng Zhang , Chenglin Li , Yin Zhang

We propose a new paradigm to automatically generate training data with accurate labels at scale using the text-to-image synthesis frameworks (e.g., DALL-E, Stable Diffusion, etc.). The proposed approach1 decouples training data generation…

计算机视觉与模式识别 · 计算机科学 2023-09-13 Yunhao Ge , Jiashu Xu , Brian Nlong Zhao , Neel Joshi , Laurent Itti , Vibhav Vineet

Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Quanjian Song , Donghao Zhou , Jingyu Lin , Fei Shen , Jiaze Wang , Xiaowei Hu , Cunjian Chen , Pheng-Ann Heng

Personalized text-to-image (P-T2I) generation aims to create new, text-guided images featuring the personalized subject with a few reference images. However, balancing the trade-off relationship between prompt fidelity and identity…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Kangyeol Kim , Wooseok Seo , Sehyun Nam , Bodam Kim , Suhyeon Jeong , Wonwoo Cho , Jaegul Choo , Youngjae Yu

Text-to-image (T2I) research has grown explosively in the past year, owing to the large-scale pre-trained diffusion models and many emerging personalization and editing approaches. Yet, one pain point persists: the text prompt engineering,…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Xingqian Xu , Jiayi Guo , Zhangyang Wang , Gao Huang , Irfan Essa , Humphrey Shi

Text-to-image (T2I) diffusion models have achieved widespread success due to their ability to generate high-resolution, photorealistic images. These models are trained on large-scale datasets, like LAION-5B, often scraped from the internet.…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Korada Sri Vardhana , Shrikrishna Lolla , Soma Biswas

Recent advances in text-to-image (T2I) diffusion models have significantly improved the quality of generated images. However, providing efficient control over individual subjects, particularly the attributes characterizing them, remains a…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Stefan Andreas Baumann , Felix Krause , Michael Neumayr , Nick Stracke , Melvin Sevi , Vincent Tao Hu , Björn Ommer

Text-to-Image (T2I) has been prevalent in recent years, with most common condition tasks having been optimized nicely. Besides, counterfactual Text-to-Image is obstructing us from a more versatile AIGC experience. For those scenes that are…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Sifan Li , Ming Tao , Hao Zhao , Ling Shao , Hao Tang

We propose a method for scene-level sketch-to-photo synthesis with text guidance. Although object-level sketch-to-photo synthesis has been widely studied, whole-scene synthesis is still challenging without reference photos that adequately…

计算机视觉与模式识别 · 计算机科学 2023-02-15 AprilPyone MaungMaung , Makoto Shing , Kentaro Mitsui , Kei Sawada , Fumio Okura

Creating content with specified identities (ID) has attracted significant interest in the field of generative models. In the field of text-to-image generation (T2I), subject-driven creation has achieved great progress with the identity…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Ze Ma , Daquan Zhou , Chun-Hsiao Yeh , Xue-She Wang , Xiuyu Li , Huanrui Yang , Zhen Dong , Kurt Keutzer , Jiashi Feng

Generating continuous sign language videos from discrete segments is challenging due to the need for smooth transitions that preserve natural flow and meaning. Traditional approaches that simply concatenate isolated signs often result in…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Shengeng Tang , Jiayi He , Lechao Cheng , Jingjing Wu , Dan Guo , Richang Hong

We focus on the foundational task of Scene Staging: given a reference scene image and a text condition specifying an actor category to be generated in the scene and its spatial relation to the scene, the goal is to synthesize an output…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Cong Xie , Che Wang , Yan Zhang , Ruiqi Yu , Han Zou , Zheng Pan , Zhenpeng Zhan

Recently, the strong latent Diffusion Probabilistic Model (DPM) has been applied to high-quality Text-to-Image (T2I) generation (e.g., Stable Diffusion), by injecting the encoded target text prompt into the gradually denoised diffusion…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Mingyang Yi , Aoxue Li , Yi Xin , Zhenguo Li

Recent advancements in text-to-image (T2I) have improved synthesis results, but challenges remain in layout control and generating omnidirectional panoramic images. Dense T2I (DT2I) and spherical T2I (ST2I) models address these issues, but…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Timon Winter , Stanislav Frolov , Brian Bernhard Moser , Andreas Dengel

3D scene generation conditioned on text prompts has significantly progressed due to the development of 2D diffusion generation models. However, the textual description of 3D scenes is inherently inaccurate and lacks fine-grained control…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Minglin Chen , Longguang Wang , Sheng Ao , Ye Zhang , Kai Xu , Yulan Guo

Person re-identification (ReID) aims to retrieve target pedestrian images given either visual queries (image-to-image, I2I) or textual descriptions (text-to-image, T2I). Although both tasks share a common retrieval objective, they pose…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Linhan Zhou , Shuang Li , Neng Dong , Yonghang Tai , Yafei Zhang , Huafeng Li

2D concept art generation for 3D scenes is a crucial yet challenging task in computer graphics, as creating natural intuitive environments still demands extensive manual effort in concept design. While generative AI has simplified 2D…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Zhenhong Sun , Yifu Wang , Yonhon Ng , Yongzhi Xu , Daoyi Dong , Hongdong Li , Pan Ji

Due to the demand for personalizing image generation, subject-driven text-to-image generation method, which creates novel renditions of an input subject based on text prompts, has received growing research interest. Existing methods often…

计算机视觉与模式识别 · 计算机科学 2025-01-08 Shang Chai , Zihang Lin , Min Zhou , Xubin Li , Liansheng Zhuang , Houqiang Li

Scene text recognition has been an important, active research topic in computer vision for years. Previous approaches mainly consider text as 1D signals and cast scene text recognition as a sequence prediction problem, by feat of CTC or…

计算机视觉与模式识别 · 计算机科学 2019-07-24 Zhaoyi Wan , Fengming Xie , Yibo Liu , Xiang Bai , Cong Yao