English

A Comprehensive Comparison of Pre-training Language Models

Computation and Language 2023-07-27 v9

Abstract

Recently, the development of pre-trained language models has brought natural language processing (NLP) tasks to the new state-of-the-art. In this paper we explore the efficiency of various pre-trained language models. We pre-train a list of transformer-based models with the same amount of text and the same training steps. The experimental results shows that the most improvement upon the origin BERT is adding the RNN-layer to capture more contextual information for short text understanding. But the conclusion is: There are no remarkable improvement for short text understanding for similar BERT structures. Data-centric method[12] can achieve better performance.

Keywords

Cite

@article{arxiv.2106.11483,
  title  = {A Comprehensive Comparison of Pre-training Language Models},
  author = {Tong Guo},
  journal= {arXiv preprint arXiv:2106.11483},
  year   = {2023}
}
R2 v1 2026-06-24T03:27:00.587Z