中文
相关论文

相关论文: Explaining CLIP through Co-Creative Drawings and I…

200 篇论文

The multimodal deep neural networks, represented by CLIP, have generated rich downstream applications owing to their excellent performance, thus making understanding the decision-making process of CLIP an essential research topic. Due to…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Chenming Shang , Hengyuan Zhang , Hao Wen , Yujiu Yang

We propose a novel framework for image clustering that incorporates joint representation learning and clustering. Our method consists of two heads that share the same backbone network - a "representation learning" head and a "clustering"…

计算机视觉与模式识别 · 计算机科学 2021-07-27 Kien Do , Truyen Tran , Svetha Venkatesh

Generative image models have emerged as a promising technology to produce realistic images. Despite potential benefits, concerns grow about its misuse, particularly in generating deceptive images that could raise significant ethical, legal,…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Jinbin Huang , Chen Chen , Aditi Mishra , Bum Chul Kwon , Zhicheng Liu , Chris Bryan

With the goal of understanding the visual concepts that CLIP associates with text prompts, we show that the latent space of CLIP can be visualized solely in terms of linear transformations on simple geometric primitives like straight lines…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Nityanand Mathur , Shyam Marjit , Abhra Chaudhuri , Anjan Dutta

Large-scale neural network models combining text and images have made incredible progress in recent years. However, it remains an open question to what extent such models encode compositional representations of the concepts over which they…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Martha Lewis , Nihal V. Nayak , Peilin Yu , Qinan Yu , Jack Merullo , Stephen H. Bach , Ellie Pavlick

CLIP has demonstrated great versatility in adapting to various downstream tasks, such as image editing and generation, visual question answering, and video understanding. However, CLIP-based applications often suffer from misunderstandings…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Zeliang Zhang , Zhuo Liu , Mingqian Feng , Chenliang Xu

Modern tools for class-agnostic image segmentation (e.g., SegmentAnything) and open-set semantic understanding (e.g., CLIP) provide unprecedented opportunities for robot perception and mapping. While traditional closed-set metric-semantic…

In the field of design patent analysis, traditional tasks such as patent classification and patent image retrieval heavily depend on the image data. However, patent images -- typically consisting of sketches with abstract and structural…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Zhu Wang , Homaira Huda Shomee , Sathya N. Ravi , Sourav Medya

Over the past years, advances in artificial intelligence (AI) have demonstrated how AI can solve many perception and generation tasks, such as image classification and text writing, yet reasoning remains a challenge. This paper introduces…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Andreas Plesner , Turlan Kuzhagaliyev , Roger Wattenhofer

Interpreting the decision logic behind effective deep convolutional neural networks (CNN) on images complements the success of deep learning models. However, the existing methods can only interpret some specific decision logic on individual…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Peter Cho-Ho Lam , Lingyang Chu , Maxim Torgonskiy , Jian Pei , Yong Zhang , Lanjun Wang

In this paper, we demonstrate that CLIP can also be adapted to downstream tasks where its vision-language alignment is suboptimally learned during pre-training on web-crawled data, all without requiring fine-tuning. We explore the case of…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Sohee Kim , Jisu Kang , Dunam Kim , Seokju Lee

This paper presents abstract art created by neural networks and broadly recognizable across various computer vision systems. The existence of abstract forms that trigger specific labels independent of neural architecture or training set…

计算机视觉与模式识别 · 计算机科学 2019-12-10 Tom White

This paper introduces CLIPSwarm, a new algorithm designed to automate the modeling of swarm drone formations based on natural language. The algorithm begins by enriching a provided word, to compose a text prompt that serves as input to an…

机器人学 · 计算机科学 2024-03-21 Pablo Pueyo , Eduardo Montijano , Ana C. Murillo , Mac Schwager

This paper presents a computational model for conceptual shifts, based on a novelty metric applied to a vector representation generated through deep learning. This model is integrated into a co-creative design system, which enables a…

人机交互 · 计算机科学 2019-06-26 Pegah Karimi , Mary Lou Maher , Nicholas Davis , Kazjon Grace

The process of painting fosters creativity and rational planning. However, existing generative AI mostly focuses on producing visually pleasant artworks, without emphasizing the painting process. We introduce a novel task, Collaborative…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Nicola Dall'Asen , Willi Menapace , Elia Peruzzo , Enver Sangineto , Yiming Wang , Elisa Ricci

Human emotions are essentially molded by lived experiences, from which we construct personalised meaning. The engagement in such meaning-making process has been practiced as an intervention in various psychotherapies to promote wellness.…

人机交互 · 计算机科学 2024-03-06 Qian Wan , Xin Feng , Yining Bei , Zhiqi Gao , Zhicong Lu

While the potential of deep learning (DL) for automating simple tasks is already well explored, recent research has started investigating the use of deep learning for creative design, both for complete artifact creation and supporting…

人工智能 · 计算机科学 2024-02-13 Marcus Basalla , Johannes Schneider , Jan vom Brocke

Verifying the authenticity of AI-generated images presents a growing challenge on social media platforms these days. While vision-language models (VLMs) like CLIP outdo in multimodal representation, their capacity for AI-generated image…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Ziyang Ou

As artificial intelligence shifts from pure tool for delegation toward agentic collaboration, its use in the arts can shift beyond the exploration of machine autonomy toward synergistic co-creation. While our earlier robotic works utilized…

人机交互 · 计算机科学 2026-03-09 Patrick Tresset , Markus Wulfmeier

The classic duck-rabbit illusion reveals that when visual evidence is ambiguous, the human brain must decide what it sees. But where exactly do human observers draw the line between ''duck'' and ''rabbit'', and do machine classifiers draw…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yuqi Hu , Vasha DuTell , Ahna R. Girshick , Jennifer E. Corbett