视觉图腾识别:精选比较数据集与分类方法的深化研究
计算机视觉与模式识别
2024-10-22 v1
摘要
在电影中,视觉图腾是指重复出现的象征性构图,承载着艺术或审美意义。其在视觉艺术和媒介史中的运用对研究者和导演而言都具有重要价值。本文旨在提出一种新的机器学习模型,通过定制数据集实现这些图腾的识别与分类。我们展示了如何利用从CLIP模型中提取的特征,通过采用合适的损失函数和浅层网络,对图像进行20种不同图腾的分类,实验结果在测试集上达到了0.91的F1分数。我们还 presented了几项消融研究,以论证所选输入特征、网络结构及超参数的合理性。
引用
@article{arxiv.2410.15866,
title = {Visual Motif Identification: Elaboration of a Curated Comparative Dataset and Classification Methods},
author = {Adam Phillips and Daniel Grandes Rodriguez and Miriam Sánchez-Manzano and Alan Salvadó and Manuel Garin and Gloria Haro and Coloma Ballester},
journal= {arXiv preprint arXiv:2410.15866},
year = {2024}
}
备注
17 pages, 11 figures, one table, to be published in the conference proceedings of ECCV 2024