涌现的能力缺失?预训练过程中的逆缩放现象
计算与语言
2024-08-27 v2
摘要
逆缩放是否仅作为模型规模的函数出现,还是也会在训练过程中发生?我们进行了一项探索性研究,考察语言模型在语言建模任务训练期间,在特定任务上的性能是否会下降(而整体性能保持较高)。我们发现 Pythia 12B(Biderman 等人,2023)在训练过程中有 8 个任务表现出性能下降。其中五个任务(TruthfulQA-MC1、TruthfulQA-MC2、Hindsight Neglect、Memo Trap 和 Pattern Match Suppression)还表现出一致的关系:尽管整体上呈现标准(正向)缩放,但较大的语言模型在训练越多时性能下降越大。这凸显了在任何时候模型被额外数据训练时,即便其整体性能提升,也需在全部相关基准上测试性能的重要性。
引用
@article{arxiv.2305.14681,
title = {Emergent inabilities? Inverse scaling over the course of pretraining},
author = {James A. Michaelov and Benjamin K. Bergen},
journal= {arXiv preprint arXiv:2305.14681},
year = {2024}
}
备注
Accepted to Findings of EMNLP 2023