语言表征空间中的低维结构反映于脑响应
计算与语言
2025-12-11 v5 机器学习
摘要
神经语言模型、翻译模型与语言标注任务所学到的表征之间有多相关?我们通过将计算机视觉中的编码器-解码器迁移学习方法加以改造来回答这一问题,以探究从各种语言任务训练的网络隐藏表征中提取的100种不同特征空间之间的结构。该方法揭示了一种低维结构,其中语言模型与翻译模型在词嵌入、句法与语义任务以及未来词嵌入之间平滑插值。我们称这一低维结构为语言表征嵌入(language representation embedding),因为它编码了处理各类NLP任务所需表征之间的关系。我们发现该表征嵌入能够预测每个独立特征空间映射到使用fMRI记录的人脑对自然语言刺激响应的程度。此外,我们发现该结构的主维度可用于构建一种度量,从而凸显大脑的自然语言处理层级。这表明该嵌入捕捉了大脑自然语言表征结构的某些部分。
引用
@article{arxiv.2106.05426,
title = {Low-Dimensional Structure in the Space of Language Representations is Reflected in Brain Responses},
author = {Richard Antonello and Javier Turek and Vy Vo and Alexander Huth},
journal= {arXiv preprint arXiv:2106.05426},
year = {2025}
}
备注
Accepted to the Advances in Neural Information Processing Systems 34 (2021) Revised to include voxel selection details