通过主动遗忘探索预训练以提高解码器语言模型的跨语言迁移
计算与语言
2025-05-22 v2
摘要
大型语言模型 (LLM) 在众多 NLP 任务中展现出卓越的能力。然而,这些模型对除英语之外的语言应用效果往往有限。已有研究表明,编码器模型如 BERT 或 XLM-RoBERTa 在从英语到其他语言方面显示出惊人的跨语言迁移能力。在本工作中,我们提出了一种预训练策略,通过主动遗忘来实现在解码器-only LLM 中实现类似的跨语言迁移。我们表明,采用主动遗忘进行预训练的 LLM 在适应新型和未见语言方面效果极佳。通过大量实验,我们发现采用主动遗忘进行预训练的 LLM 能够学习更好的多语言表征,这在诸多下游任务中 translates to better performance。
引用
@article{arxiv.2410.16168,
title = {Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models},
author = {Divyanshu Aggarwal and Ashutosh Sathe and Sunayana Sitaram},
journal= {arXiv preprint arXiv:2410.16168},
year = {2025}
}
备注
12 pages, 11 tables, 12 figures