中文
相关论文

相关论文: GraPE: A Generate-Plan-Edit Framework for Composit…

200 篇论文

Recent strides in Text-to-3D techniques have been propelled by distilling knowledge from powerful large text-to-image diffusion models (LDMs). Nonetheless, existing Text-to-3D approaches often grapple with challenges such as…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Yiwen Chen , Chi Zhang , Xiaofeng Yang , Zhongang Cai , Gang Yu , Lei Yang , Guosheng Lin

Text-to-Image (T2I) generation has made significant advancements with diffusion models, yet challenges persist in handling complex instructions, ensuring fine-grained content control, and maintaining deep semantic consistency. Existing T2I…

机器学习 · 计算机科学 2025-08-08 Xiaoqi Dong , Xiangyu Zhou , Nicholas Evans , Yujia Lin

Text-to-Image (T2I) models have demonstrated impressive capabilities in generating high-quality and diverse visual content from natural language prompts. However, uncontrolled reproduction of sensitive, copyrighted, or harmful imagery poses…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Yiwei Xie , Ping Liu , Zheng Zhang

To replicate the success of text-to-image (T2I) generation, recent works employ large-scale video datasets to train a text-to-video (T2V) generator. Despite their promising results, such paradigm is computationally expensive. In this work,…

计算机视觉与模式识别 · 计算机科学 2023-03-20 Jay Zhangjie Wu , Yixiao Ge , Xintao Wang , Weixian Lei , Yuchao Gu , Yufei Shi , Wynne Hsu , Ying Shan , Xiaohu Qie , Mike Zheng Shou

Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scenes, often collapsing to repetitive layouts, stereotypical poses, and poorly grounded…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Wenxuan Peng , Bharath Hariharan , Hadar Averbuch-Elor

Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models optimize for "average" human appeal, they fail to capture the inherent subjectivity of…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Anne-Sofie Maerten , Juliane Verwiebe , Shyamgopal Karthik , Ameya Prabhu , Johan Wagemans , Matthias Bethge

Pre-trained large text-to-image models synthesize impressive images with an appropriate use of text prompts. However, ambiguities inherent in natural language and out-of-distribution effects make it hard to synthesize image styles, that…

Making text-to-image (T2I) generative model sample both fast and well represents a promising research direction. Previous studies have typically focused on either enhancing the visual quality of synthesized images at the expense of sampling…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Shitong Shao , Zikai Zhou , Dian Xie , Yuetong Fang , Tian Ye , Lichen Bai , Zeke Xie

Recent advances in text-to-image (T2I) models, especially diffusion-based architectures, have significantly improved the visual quality of generated images. However, these models continue to struggle with a critical limitation: maintaining…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Yifan Shen , Yangyang Shu , Hye-young Paik , Yulei Sui

Text-to-image (T2I) generation model has made significant advancements, resulting in high-quality images aligned with an input prompt. However, despite T2I generation's ability to generate fine-grained images, it still faces challenges in…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Taekyung Lee , Donggyu Lee , Myungjoo Kang

Recent advances in text-guided image synthesis has dramatically changed how creative professionals generate artistic and aesthetically pleasing visual assets. To fully support such creative endeavors, the process should possess the ability…

计算机视觉与模式识别 · 计算机科学 2023-10-31 K J Joseph , Prateksha Udhayanan , Tripti Shukla , Aishwarya Agarwal , Srikrishna Karanam , Koustava Goswami , Balaji Vasan Srinivasan

Although text-to-image generation technologies have made significant advancements, they still face challenges when dealing with ambiguous prompts and aligning outputs with user intent.Our proposed framework, TDRI (Two-Phase Dialogue…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Yuheng Feng , Jianhui Wang , Kun Li , Sida Li , Tianyu Shi , Haoyue Han , Miao Zhang , Xueqian Wang

Diffusion models have recently achieved remarkable advancements in terms of image quality and fidelity to textual prompts. Concurrently, the safety of such generative models has become an area of growing concern. This work introduces a…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Tong Liu , Zhixin Lai , Jiawen Wang , Gengyuan Zhang , Shuo Chen , Philip Torr , Vera Demberg , Volker Tresp , Jindong Gu

Scalable Vector Graphics (SVGs) are highly favored by designers due to their resolution independence and well-organized layer structure. Although existing text-to-vector (T2V) generation methods can create SVGs from text prompts, they often…

图形学 · 计算机科学 2025-05-16 Peiying Zhang , Nanxuan Zhao , Jing Liao

Text-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of…

人机交互 · 计算机科学 2023-09-29 Vivian Liu , Lydia B. Chilton

Text-to-image (T2I) generation using multiple conditions enables fine-grained user control on the generated image. Yet, incorporating multi-condition inputs incurs substantial computation and communication overhead, due to additional…

多媒体 · 计算机科学 2026-05-12 Yuxin Kong , Peng Yang , Chongbin Yi , Fan Wu , Feng Lyu

Generating images from textual descriptions has recently attracted a lot of interest. While current models can generate photo-realistic images of individual objects such as birds and human faces, synthesising images with multiple objects is…

计算机视觉与模式识别 · 计算机科学 2020-10-29 Stanislav Frolov , Shailza Jolly , Jörn Hees , Andreas Dengel

Text-to-image (T2I) models are capable of generating visually impressive images, yet they often fail to accurately capture specific attributes in user prompts, such as the correct number of objects with the specified colors. The diversity…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Kevin David Hayes , Micah Goldblum , Vikash Sehwag , Gowthami Somepalli , Ashwinee Panda , Tom Goldstein

Recently, Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated metrics have emerged to evaluate the image-text alignment capabilities of generative models. However, the performance…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Shuhao Han , Haotian Fan , Jiachen Fu , Liang Li , Tao Li , Junhui Cui , Yunqiu Wang , Yang Tai , Jingwei Sun , Chunle Guo , Chongyi Li

We address the problem of interactive text-to-image (T2I) generation, designing a reinforcement learning (RL) agent which iteratively improves a set of generated images for a user through a sequence of prompt expansions. Using human raters,…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Ofir Nabati , Guy Tennenholtz , ChihWei Hsu , Moonkyung Ryu , Deepak Ramachandran , Yinlam Chow , Xiang Li , Craig Boutilier
‹ 上一页 1 8 9 10 下一页 ›