什么不在何处:将空间表示集成到深度学习架构中的挑战
机器学习
2018-07-24 v1 人工智能
计算与语言
神经与进化计算
机器学习
摘要
本文考察当前用于图像描述生成的深度学习架构在多大程度上捕捉了空间语言。基于对文献中生成描述示例的评估,我们认为这些系统捕捉了图像数据中有哪些对象,但未捕捉这些对象位于何处:这些系统生成的描述是以对象检测器输出为条件的语言模型的输出,而该检测器无法捕捉细粒度位置信息。尽管语言模型为图像描述提供了有用的知识,我们认为深度学习图像描述架构也应当建模对象间的几何关系。
引用
@article{arxiv.1807.08133,
title = {What is not where: the challenge of integrating spatial representations into deep learning architectures},
author = {John D. Kelleher and Simon Dobnik},
journal= {arXiv preprint arXiv:1807.08133},
year = {2018}
}
备注
15 pages, 10 figures, Appears in CLASP Papers in Computational Linguistics Vol 1: Proceedings of the Conference on Logic and Machine Learning in Natural Language (LaML 2017), pp. 41-52