English
Related papers

Related papers: On Semiotic-Grounded Interpretive Evaluation of Ge…

200 papers

Effectively aligning with human judgment when evaluating machine-generated image captions represents a complex yet intriguing challenge. Existing evaluation metrics like CIDEr or CLIP-Score fall short in this regard as they do not take into…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Sara Sarto , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

With recent advances in image-to-image translation tasks, remarkable progress has been witnessed in generating face images from sketches. However, existing methods frequently fail to generate images with details that are semantically and…

Computer Vision and Pattern Recognition · Computer Science 2023-02-09 Binxin Yang , Xuejin Chen , Chaoqun Wang , Chi Zhang , Zihan Chen , Xiaoyan Sun

Memes represent a tightly coupled, multimodal form of social expression, in which visual context and overlaid text jointly convey nuanced affect and commentary. Inspired by cognitive reappraisal in psychology, we introduce Meme Reappraisal,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yiqi Nie , Fei Wang , Junjie Chen , Kun Li , Yudi Cai , Dan Guo , Chenglong Li , Meng Wang

Semantic image editing requires inpainting pixels following a semantic map. It is a challenging task since this inpainting requires both harmony with the context and strict compliance with the semantic maps. The majority of the previous…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Hakan Sivuk , Aysegul Dundar

Semantic image synthesis aims to generate high-quality images given semantic conditions, i.e. segmentation masks and style reference images. Existing methods widely adopt generative adversarial networks (GANs). GANs take all conditional…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Feng Liu , Xiaobin Chang

The current multimodal turn in linguistic theory leaves a crucial question unanswered: what is the meaning of iconic gestures, and how does it compose with speech meaning? We argue for a separation of linguistic and visual levels of meaning…

Computation and Language · Computer Science 2025-12-12 Andy Lücking , Alexander Henlein , Alexander Mehler

Allowing users to interact through language borders is an interesting challenge for information technology. For the purpose of a computer assisted language learning system, we have chosen icons for representing meaning on the input…

Computation and Language · Computer Science 2007-05-23 Pascal Vaillant

Key doctrines, including novelty (patent), originality (copyright), and distinctiveness (trademark), turn on a shared empirical question: whether a body of work is meaningfully distinct from a relevant reference class. Yet analyses…

Computers and Society · Computer Science 2026-01-27 Anirban Mukherjee , Hannah Hanwen Chang

Semantic communication has recently attracted significant interest from both industry and academia due to its potential to transform the existing data-focused communication architecture towards a more generally intelligent and goal-oriented…

Artificial Intelligence · Computer Science 2023-01-16 Yong Xiao , Zijian Sun , Guangming Shi , Dusit Niyato

Semantic image synthesis (SIS) refers to the problem of generating realistic imagery given a semantic segmentation mask that defines the spatial layout of object classes. Most of the approaches in the literature, other than the quality of…

Computer Vision and Pattern Recognition · Computer Science 2023-07-12 Tomaso Fontanini , Claudio Ferrari , Massimo Bertozzi , Andrea Prati

When retrieval-augmented generation (RAG) systems hallucinate, what geometric trace does this leave in embedding space? We introduce the Semantic Grounding Index (SGI), defined as the ratio of angular distances from the response to the…

Artificial Intelligence · Computer Science 2025-12-17 Javier Marín

Analyzing digitized artworks presents unique challenges, requiring not only visual interpretation but also a deep understanding of rich artistic, contextual, and historical knowledge. We introduce ArtSeek, a multimodal framework for art…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Nicola Fanelli , Gennaro Vessio , Giovanna Castellano

Implicit artistic influence, although visually plausible, is often undocumented and thus poses a historically constrained attribution problem: resemblance is necessary but not sufficient evidence. Most prior systems reduce influence…

Artificial Intelligence · Computer Science 2026-04-10 Hanyi Liu , Zhonghao Jiu , Minghao Wang , Yuhang Xie , Heran Yang

Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over explicit cultural context, making it difficult to jointly maintain image relevance,…

Computation and Language · Computer Science 2026-04-21 Run Xu , Lu Li , Rongzhao Zhang , Jie Xu

With the advancement of neural generative capabilities, the art community has increasingly embraced GenAI (Generative Artificial Intelligence), particularly large text-to-image models, for producing aesthetically compelling results.…

Human-Computer Interaction · Computer Science 2025-08-26 Aven-Le Zhou , Wei Wu , Yu-Ao Wang , Kang Zhang

CLIP has emerged as a powerful multimodal model capable of connecting images and text through joint embeddings, but to what extent does it 'see' the same way humans do - especially when interpreting artworks? In this paper, we investigate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Andrea Asperti , Leonardo Dessì , Maria Chiara Tonetti , Nico Wu

It has long been hypothesized that perceptual ambiguities play an important role in aesthetic experience: a work with some ambiguity engages a viewer more than one that does not. However, current frameworks for testing this theory are…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Xi Wang , Zoya Bylinskii , Aaron Hertzmann , Robert Pepperell

This paper proposes new framework of communication system leveraging promising generation capabilities of multi-modal generative models. Regarding nowadays smart applications, successful communication can be made by conveying the perceptual…

Signal Processing · Electrical Eng. & Systems 2023-09-11 Hyelin Nam , Jihong Park , Jinho Choi , Seong-Lyun Kim

We propose SemCSE-Multi, a novel unsupervised framework for generating multifaceted embeddings of scientific abstracts, evaluated in the domains of invasion biology and medicine. These embeddings capture distinct, individually specifiable…

Computation and Language · Computer Science 2026-01-13 Marc Brinner , Sina Zarrieß

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual understanding, they often struggle when faced with the unstructured and ambiguous nature of human-generated sketches. This limitation is particularly…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Yuhang Su , Mei Wang , Yaoyao Zhong , Guozhang Li , Shixing Li , Yihan Feng , Hua Huang