English

A Survey of Pre-trained Language Models for Processing Scientific Text

Computation and Language 2024-02-01 v1

Abstract

The number of Language Models (LMs) dedicated to processing scientific text is on the rise. Keeping pace with the rapid growth of scientific LMs (SciLMs) has become a daunting task for researchers. To date, no comprehensive surveys on SciLMs have been undertaken, leaving this issue unaddressed. Given the constant stream of new SciLMs, appraising the state-of-the-art and how they compare to each other remain largely unknown. This work fills that gap and provides a comprehensive review of SciLMs, including an extensive analysis of their effectiveness across different domains, tasks and datasets, and a discussion on the challenges that lie ahead.

Keywords

Cite

@article{arxiv.2401.17824,
  title  = {A Survey of Pre-trained Language Models for Processing Scientific Text},
  author = {Xanh Ho and Anh Khoa Duong Nguyen and An Tuan Dao and Junfeng Jiang and Yuki Chida and Kaito Sugimoto and Huy Quoc To and Florian Boudin and Akiko Aizawa},
  journal= {arXiv preprint arXiv:2401.17824},
  year   = {2024}
}

Comments

Resources are available at https://github.com/Alab-NII/Awesome-SciLM

R2 v1 2026-06-28T14:33:02.919Z