中文

Transformer中统计偏差存在下的泛化与记忆

机器学习 2024-09-11 v1 机器学习

摘要

本研究旨在理解统计偏差如何影响模型在算法任务上对分布内和分布外数据的泛化能力。先前的研究表明,Transformer可能无意中学会依赖这些虚假相关性,导致对其泛化能力的高估。为了研究这一点,我们在几个合成的算法任务上评估了Transformer模型,系统地引入并改变了这些偏差的存在。我们还分析了Transformer模型的不同组件如何影响其泛化能力。我们的发现表明,统计偏差损害了模型在分布外数据上的性能,导致对其泛化能力的高估。模型严重依赖这些虚假相关性进行推理,这从它们在包含此类偏差的任务上的表现可以看出。

关键词

引用

@article{arxiv.2409.04654,
  title  = {Generalization vs. Memorization in the Presence of Statistical Biases in Transformers},
  author = {John Mitros},
  journal= {arXiv preprint arXiv:2409.04654},
  year   = {2024}
}

备注

arXiv admin note: This submission has been removed by arXiv administrators as the user did not have the right to agree to arXiv's license at the time of the submission. The author list has been truncated