fruit-SALAD:一种风格对齐艺术品数据集以揭示图像嵌入中的相似性感知
计算机视觉与模式识别
2025-11-12 v1 人工智能
计算复杂性
机器学习
摘要
视觉相似性概念对于 computer vision 以及围绕图像向量嵌入的应用和研究至关重要。然而,稀缺的基准数据集是一个在探索这些模型如何感知相似性方面具有重大挑战的障碍。本文我们引入 Style Aligned Artwork Datasets(SALADs),并以 fruit-SALAD 为例,包含 10,000 张水果描绘图像。该 combined semantic category 和 style benchmark 包含 10 个易于识别的水果类别中每组 100 个实例,跨 10 个易于区分的风格。利用 system pipeline of generative image synthesis,此 visually diverse yet balanced benchmark 在 various computational models 中展示了显著差异,包括 machine learning models、feature extraction algorithms、complexity measures,以及用于 reference 的概念模型。这个 meticulously 设计的数据集为 comparative analysis of similarity perception 提供了一个受控且平衡的平台。SALAD 框架允许比较这些模型在 semantic category 和 style recognition 任务中的性能,超越 anecdotal knowledge 的水平,使其变得 robustly quantifiable and qualitatively interpretable。
引用
@article{arxiv.2406.01278,
title = {fruit-SALAD: A Style Aligned Artwork Dataset to reveal similarity perception in image embeddings},
author = {Tillmann Ohm and Andres Karjus and Mikhail Tamm and Maximilian Schich},
journal= {arXiv preprint arXiv:2406.01278},
year = {2025}
}