VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
Abstract
A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipulate concept-level properties under transformation and transfer them to new scenes. We introduce VisAnalog, a controlled suite for this setting on natural images. Each example instantiates : images and a hidden target image are produced by applying the same deterministic transformation sequence to source images and . Given , , and , a model must answer a multiple-choice question about . The benchmark contains 617 human-validated questions spanning one- to four-step transformations such as zoom, quadrant swap, rotation, flip, and hue rotation. Across strong proprietary and open-source VLMs, end-to-end accuracy is substantially lower than oracle accuracy when is directly shown, and degrades sharply as transformation depth increases, while human performance remains near the ceiling. A program-conditioned evaluation further separates failures of relation inference from failures of transformation application, showing that inferring the visual relation from is the dominant bottleneck, with additional application errors emerging on harder multi-step cases. The dataset is publicly available at https://huggingface.co/datasets/zli99/VisAnalog.
Cite
@article{arxiv.2605.23141,
title = {VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images},
author = {Zhaonan Li and Kyle R. Chickering and Bangzheng Li and Jacob Dineen and Xiao Ye and Zhikun Xu and Shijie Lu and Yuxi Huang and Ming Shen and Bach Nguyen and Jaya Adithya Pavuluri and Mau Son Nguyen and Sanika Chavan and Ngoc Minh Thu Le and Muhao Chen and Ben Zhou},
journal= {arXiv preprint arXiv:2605.23141},
year = {2026}
}
Comments
Accepted to the Workshop on Visual Concepts at CVPR 2026 as a non-archival report