语言是否有助于视觉模型的泛化?
人工智能
2021-09-16 v3 计算与语言
计算机视觉与模式识别
摘要
在大规模多模态数据集上训练的视觉模型可受益于广泛可用的图像- caption 数据集。近期一个模型(CLIP)被发现能在零样本和迁移学习设定中良好泛化。这可能意味着语言或“语义接地”赋予视觉特征空间额外的泛化能力。在此,我们系统评估各类多模态架构与纯视觉模型在无监督聚类、少样本学习、迁移学习及对抗鲁棒性方面的表现。在每种设定中,相比标准监督视觉训练,多模态训练并未产生额外的泛化能力。我们得出结论:要使语义接地有助于改进视觉模型,仍需开展进一步工作。
引用
@article{arxiv.2104.08313,
title = {Does language help generalization in vision models?},
author = {Benjamin Devillers and Bhavin Choksi and Romain Bielawski and Rufin VanRullen},
journal= {arXiv preprint arXiv:2104.08313},
year = {2021}
}
备注
Paper accepted at the CoNLL 2021 conference. This version: section added on the performance of the visual and visio-linguistic models on linquistic tasks