中文
相关论文

相关论文: Composition and Deformance: Measuring Imageability…

200 篇论文

The task of text-to-image generation has encountered significant challenges when applied to literary works, especially poetry. Poems are a distinct form of literature, with meanings that frequently transcend beyond the literal words. To…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Sofia Jamil , Bollampalli Areen Reddy , Raghvendra Kumar , Sriparna Saha , K J Joseph , Koustava Goswami

The notions of concreteness and imageability, traditionally important in psycholinguistics, are gaining significance in semantic-oriented natural language processing tasks. In this paper we investigate the predictability of these two…

计算与语言 · 计算机科学 2022-09-15 Nikola Ljubešić , Darja Fišer , Anita Peti-Stantić

Research in Image Generation has recently made significant progress, particularly boosted by the introduction of Vision-Language models which are able to produce high-quality visual content based on textual inputs. Despite ongoing…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Federico Betti , Jacopo Staiano , Lorenzo Baraldi , Lorenzo Baraldi , Rita Cucchiara , Nicu Sebe

Subject-Driven Text-to-Image (T2I) Generation aims to preserve a subject's identity while editing its context based on a text prompt. A core challenge in this task is the "similarity-controllability paradox", where enhancing textual control…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Shuang Li , Chao Deng , Hang Chen , Liqun Liu , Zhenyu Hu , Te Cao , Mengge Xue , Yuan Chen , Peng Shu , Huan Yu , Jie Jiang

We introduce GRADE, an automatic method for quantifying sample diversity in text-to-image models. Our method leverages the world knowledge embedded in large language models and visual question-answering systems to identify relevant…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Royi Rassin , Aviv Slobodkin , Shauli Ravfogel , Yanai Elazar , Yoav Goldberg

Existing automatic evaluation on text-to-image synthesis can only provide an image-text matching score, without considering the object-level compositionality, which results in poor correlation with human judgments. In this work, we propose…

计算机视觉与模式识别 · 计算机科学 2023-05-19 Yujie Lu , Xianjun Yang , Xiujun Li , Xin Eric Wang , William Yang Wang

Words can be represented by composing the representations of subword units such as word segments, characters, and/or character n-grams. While such representations are effective and may capture the morphological regularities of words, they…

计算与语言 · 计算机科学 2017-04-28 Clara Vania , Adam Lopez

Conditional image generation is an active research topic including text2image and image translation. Recently image manipulation with linguistic instruction brings new challenges of multimodal conditional generation. However, traditional…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Zhenhuan Liu , Jincan Deng , Liang Li , Shaofei Cai , Qianqian Xu , Shuhui Wang , Qingming Huang

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different from the assessment of text-driven image generation,…

计算机视觉与模式识别 · 计算机科学 2025-01-20 Shangkun Sun , Bowen Qu , Xiaoyu Liang , Songlin Fan , Wei Gao

Most state-of-the-art image retrieval and recommendation systems predominantly focus on individual images. In contrast, socially curated image collections, condensing distinctive yet coherent images into one set, are largely overlooked by…

多媒体 · 计算机科学 2016-11-17 Yuncheng Li , Yang Cong , Tao Mei , Jiebo Luo

Relations are basic building blocks of human cognition. Classic and recent work suggests that many relations are early developing, and quickly perceived. Machine models that aspire to human-level perception and reasoning should reflect the…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Colin Conwell , Tomer Ullman

Image-to-image translation tasks have been widely investigated with Generative Adversarial Networks (GANs) and dual learning. However, existing models lack the ability to control the translated results in the target domain and their results…

计算机视觉与模式识别 · 计算机科学 2018-05-02 Jianxin Lin , Yingce Xia , Tao Qin , Zhibo Chen , Tie-Yan Liu

The transformative potential of text-to-image (T2I) models hinges on their ability to synthesize culturally diverse, photorealistic images from textual prompts. However, these models often perpetuate cultural biases embedded within their…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Muna Numan Said , Aarib Zaidi , Rabia Usman , Sonia Okon , Praneeth Medepalli , Kevin Zhu , Vasu Sharma , Sean O'Brien

In the creative practice of text-to-image (TTI) generation, images are synthesized from textual prompts. By design, TTI models always yield an output, even if the prompt contains unknown terms. In this case, the model may generate default…

人机交互 · 计算机科学 2026-01-27 Hannu Simonen , Atte Kiviniemi , Hannah Johnston , Helena Barranha , Jonas Oppenlaender

State-of-the-art T2I models are capable of generating high-resolution images given textual prompts. However, they still struggle with accurately depicting compositional scenes that specify multiple objects, attributes, and spatial…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Yixin Wan , Kai-Wei Chang

State-of-the-art Text-to-Image models like Stable Diffusion and DALLE$\cdot$2 are revolutionizing how people generate visual content. At the same time, society has serious concerns about how adversaries can exploit such models to generate…

计算机视觉与模式识别 · 计算机科学 2023-08-17 Yiting Qu , Xinyue Shen , Xinlei He , Michael Backes , Savvas Zannettou , Yang Zhang

Text-to-image diffusion has attracted vast attention due to its impressive image-generation capabilities. However, when it comes to human-centric text-to-image generation, particularly in the context of faces and hands, the results often…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Jie Zhu , Yixiong Chen , Mingyu Ding , Ping Luo , Leye Wang , Jingdong Wang

Recent advancements in diffusion models have significantly impacted the trajectory of generative machine learning research, with many adopting the strategy of fine-tuning pre-trained models using domain-specific text-to-image datasets.…

计算机视觉与模式识别 · 计算机科学 2023-12-21 Mischa Dombrowski , Hadrien Reynaud , Johanna P. Müller , Matthew Baugh , Bernhard Kainz

Word embeddings are usually derived from corpora containing text from many individuals, thus leading to general purpose representations rather than individually personalized representations. While personalized embeddings can be useful to…

计算与语言 · 计算机科学 2020-11-22 Charles Welch , Jonathan K. Kummerfeld , Verónica Pérez-Rosas , Rada Mihalcea

The architectural blueprint of today's leading text-to-image models contains a fundamental flaw: an inability to handle logical composition. This survey investigates this breakdown across three core primitives-negation, counting, and…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Mayank Vatsa , Aparna Bharati , Richa Singh