单个向量中能塞进什么:探查句子嵌入的语言学性质
计算与语言
2018-07-10 v2
摘要
尽管近来大量工作致力于训练高质量的句子嵌入,我们对其所捕获的内容仍知之甚少。基于句子分类的“下游”任务常被用来评估句子表示的质量。然而,任务的复杂性使得我们难以推断表示中究竟包含何种信息。本文引入 10 个探查任务,旨在捕捉句子的简单语言学特征,并用它们研究由三种不同编码器以八种不同方式训练所生成的嵌入,揭示了编码器与训练方法中若干引人入胜的性质。
引用
@article{arxiv.1805.01070,
title = {What you can cram into a single vector: Probing sentence embeddings for linguistic properties},
author = {Alexis Conneau and German Kruszewski and Guillaume Lample and Loïc Barrault and Marco Baroni},
journal= {arXiv preprint arXiv:1805.01070},
year = {2018}
}
备注
ACL 2018