English

SsciBERT: A Pre-trained Language Model for Social Science Texts

Computation and Language 2022-11-28 v3

Abstract

The academic literature of social sciences records human civilization and studies human social problems. With its large-scale growth, the ways to quickly find existing research on relevant issues have become an urgent demand for researchers. Previous studies, such as SciBERT, have shown that pre-training using domain-specific texts can improve the performance of natural language processing tasks. However, the pre-trained language model for social sciences is not available so far. In light of this, the present research proposes a pre-trained model based on the abstracts published in the Social Science Citation Index (SSCI) journals. The models, which are available on GitHub (https://github.com/S-T-Full-Text-Knowledge-Mining/SSCI-BERT), show excellent performance on discipline classification, abstract structure-function recognition, and named entity recognition tasks with the social sciences literature.

Keywords

Cite

@article{arxiv.2206.04510,
  title  = {SsciBERT: A Pre-trained Language Model for Social Science Texts},
  author = {Si Shen and Jiangfeng Liu and Litao Lin and Ying Huang and Lin Zhang and Chang Liu and Yutong Feng and Dongbo Wang},
  journal= {arXiv preprint arXiv:2206.04510},
  year   = {2022}
}

Comments

24 pages,2 figures

R2 v1 2026-06-24T11:45:09.196Z