中文

语言是否有助于视觉模型的泛化?

人工智能 2021-09-16 v3 计算与语言 计算机视觉与模式识别

摘要

在大规模多模态数据集上训练的视觉模型可受益于广泛可用的图像- caption 数据集。近期一个模型(CLIP)被发现能在零样本和迁移学习设定中良好泛化。这可能意味着语言或“语义接地”赋予视觉特征空间额外的泛化能力。在此,我们系统评估各类多模态架构与纯视觉模型在无监督聚类、少样本学习、迁移学习及对抗鲁棒性方面的表现。在每种设定中,相比标准监督视觉训练,多模态训练并未产生额外的泛化能力。我们得出结论:要使语义接地有助于改进视觉模型,仍需开展进一步工作。

关键词

引用

@article{arxiv.2104.08313,
  title  = {Does language help generalization in vision models?},
  author = {Benjamin Devillers and Bhavin Choksi and Romain Bielawski and Rufin VanRullen},
  journal= {arXiv preprint arXiv:2104.08313},
  year   = {2021}
}

备注

Paper accepted at the CoNLL 2021 conference. This version: section added on the performance of the visual and visio-linguistic models on linquistic tasks