English
Related papers

Related papers: Editor's Choice: Evaluating Abstract Intent in Ima…

200 papers

Existing text-guided image manipulation methods aim to modify the appearance of the image or to edit a few objects in a virtual or simple scenario, which is far from practical applications. In this work, we study a novel task on text-guided…

Computer Vision and Pattern Recognition · Computer Science 2023-02-23 Yikai Wang , Jianan Wang , Guansong Lu , Hang Xu , Zhenguo Li , Wei Zhang , Yanwei Fu

Do large language models (LLMs) genuinely understand abstract concepts, or merely manipulate them as statistical patterns? We introduce an abstraction-grounding framework that decomposes conceptual understanding into three capacities:…

Computation and Language · Computer Science 2026-01-21 Junyu Zhang , Yipeng Kang , Jiong Guo , Jiayu Zhan , Junqi Wang

Model editing aims to correct errors in large, pretrained models without altering unrelated behaviors. While some recent works have edited vision-language models (VLMs), no existing editors tackle reasoning-heavy tasks, which typically…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Jiaxing Qiu , Kaihua Hou , Roxana Daneshjou , Ahmed Alaa , Thomas Hartvigsen

In daily life, images as common affective stimuli have widespread applications. Despite significant progress in text-driven image editing, there is limited work focusing on understanding users' emotional requests. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Peixuan Zhang , Shuchen Weng , Chengxuan Zhu , Binghao Tang , Zijian Jia , Si Li , Boxin Shi

Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly improved caption…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Israa A. Albadarneh , Bassam H. Hammo , Omar S. Al-Kadi

The principle of abstraction guides the design of interactive systems, yet we lack a conceptual framework to understand how it shapes interaction design. Existing models, such as the gulfs of execution and evaluation, do not explicitly…

Human-Computer Interaction · Computer Science 2026-05-13 Bryan Min , Sangho Suh , Jim Hollan , Haijun Xia

Text-driven image editing has achieved remarkable success in following single instructions. However, real-world scenarios often involve complex, multi-step instructions, particularly ``chain'' instructions where operations are…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Chenglin Wang , Yucheng Zhou , Qianning Wang , Zhe Wang , Kai Zhang

Studies have been conducted to prevent specific concepts from being generated from pretrained text-to-image generative models, achieving concept erasure in various ways. However, the performance evaluation of these studies is still largely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Masane Fuchi , Tomohiro Takagi

Text-to-image diffusion models have emerged as powerful tools for high-quality image generation and editing. Many existing approaches rely on text prompts as editing guidance. However, these methods are constrained by the need for manual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Yuanyuan Chang , Yinghua Yao , Tao Qin , Mengmeng Wang , Ivor Tsang , Guang Dai

Large Language Models have seen expanding application across domains, yet their effectiveness as assistive tools for scientific writing - an endeavor requiring precision, multimodal synthesis, and domain expertise - remains insufficiently…

Human-Computer Interaction · Computer Science 2026-01-28 Sanchaita Hazra , Doeun Lee , Bodhisattwa Prasad Majumder , Sachin Kumar

Entity state tracking is a necessary component of world modeling that requires maintaining coherent representations of entities over time. Previous work has benchmarked entity tracking performance in purely text-based tasks. We introduce…

Computation and Language · Computer Science 2026-02-10 Vanya Cohen , Raymond Mooney

Text-to-image (T2I) systems increasingly rely on upstream prompters, either humans or multimodal large language models (MLLMs), to translate user intent into detailed prompts. Yet current benchmarks fix the prompt and only evaluate T2I…

Artificial Intelligence · Computer Science 2026-05-22 Hanjun Luo , Zhimu Huang , Sylvia Chung , Yiran Wang , Yingbin Jin , Jialin Li , Jiang Li , Xinfeng Li , Hanan Salam

Editing images via instruction provides a natural way to generate interactive content, but it is a big challenge due to the higher requirement of scene understanding and generation. Prior work utilizes a chain of large language models,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Liya Ji , Chenyang Qi , Qifeng Chen

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different from the assessment of text-driven image generation,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-20 Shangkun Sun , Bowen Qu , Xiaoyu Liang , Songlin Fan , Wei Gao

Abstraction is a desirable capability for deep learning models, which means to induce abstract concepts from concrete instances and flexibly apply them beyond the learning context. At the same time, there is a lack of clear understanding…

Machine Learning · Computer Science 2023-02-24 Shengnan An , Zeqi Lin , Bei Chen , Qiang Fu , Nanning Zheng , Jian-Guang Lou

Beyond conventional paradigms of translating speech and text, recently, there has been interest in automated transcreation of images to facilitate localization of visual content across different cultures. Attempts to define this as a formal…

Computation and Language · Computer Science 2025-03-24 Simran Khanuja , Vivek Iyer , Claire He , Graham Neubig

Recent advances in text-to-image generation have improved the quality of synthesized images, but evaluations mainly focus on aesthetics or alignment with text prompts. Thus, it remains unclear whether these models can accurately represent a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Hsin-Ping Huang , Xinyi Wang , Yonatan Bitton , Hagai Taitelbaum , Gaurav Singh Tomar , Ming-Wei Chang , Xuhui Jia , Kelvin C. K. Chan , Hexiang Hu , Yu-Chuan Su , Ming-Hsuan Yang

Text-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensively evaluate a range of text-to-image models on numerical…

Machine Learning · Computer Science 2025-02-07 Ivana Kajić , Olivia Wiles , Isabela Albuquerque , Matthias Bauer , Su Wang , Jordi Pont-Tuset , Aida Nematzadeh

While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While previous work has evaluated T2I alignment by proposing metrics, benchmarks, and templates for…

In creativity support and computational co-creativity contexts, the task of discovering appropriate prompts for use with text-to-image generative models remains difficult. In many cases the creator wishes to evoke a certain impression with…

Artificial Intelligence · Computer Science 2023-02-21 Francisco Ibarrola , Rohan Lulham , Kazjon Grace