English

Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

Computation and Language 2020-05-07 v3 Machine Learning

Abstract

Language models pretrained on text from a wide variety of sources form the foundation of today's NLP. In light of the success of these broad-coverage models, we investigate whether it is still helpful to tailor a pretrained model to the domain of a target task. We present a study across four domains (biomedical and computer science publications, news, and reviews) and eight classification tasks, showing that a second phase of pretraining in-domain (domain-adaptive pretraining) leads to performance gains, under both high- and low-resource settings. Moreover, adapting to the task's unlabeled data (task-adaptive pretraining) improves performance even after domain-adaptive pretraining. Finally, we show that adapting to a task corpus augmented using simple data selection strategies is an effective alternative, especially when resources for domain-adaptive pretraining might be unavailable. Overall, we consistently find that multi-phase adaptive pretraining offers large gains in task performance.

Keywords

Cite

@article{arxiv.2004.10964,
  title  = {Don't Stop Pretraining: Adapt Language Models to Domains and Tasks},
  author = {Suchin Gururangan and Ana Marasović and Swabha Swayamdipta and Kyle Lo and Iz Beltagy and Doug Downey and Noah A. Smith},
  journal= {arXiv preprint arXiv:2004.10964},
  year   = {2020}
}

Comments

ACL 2020

R2 v1 2026-06-23T15:02:39.633Z