利用基于 CLIP 的多模态方法实现艺术品分类与检索
计算机视觉与模式识别
2023-09-22 v1
摘要
鉴于多模态图像预训练的最新进展——其中使用语义密集的文本监督训练的视觉模型往往比使用类别属性或通过无监督技术训练的模型具有更好的泛化能力——在本文中,我们研究了近期 CLIP 模型如何应用于艺术品领域的若干任务。我们在 NoisyArt 数据集上进行了详尽实验,该数据集是从网络公共资源爬取的艺术品图像数据集。在此数据集上,CLIP 在(零样本)分类上取得了令人印象深刻的成果,并在艺术品到艺术品以及描述到艺术品检索中都获得了有前景的结果。
引用
@article{arxiv.2309.12110,
title = {Exploiting CLIP-based Multi-modal Approach for Artwork Classification and Retrieval},
author = {Alberto Baldrati and Marco Bertini and Tiberio Uricchio and Alberto Del Bimbo},
journal= {arXiv preprint arXiv:2309.12110},
year = {2023}
}
备注
Proc. of Florence Heri-Tech 2022: The Future of Heritage Science and Technologies: ICT and Digital Heritage, 2022