English

Sublanguage: A Serious Issue Affects Pretrained Models in Legal Domain

Computation and Language 2021-09-06 v2 Artificial Intelligence

Abstract

Legal English is a sublanguage that is important for everyone but not for everyone to understand. Pretrained models have become best practices among current deep learning approaches for different problems. It would be a waste or even a danger if these models were applied in practice without knowledge of the sublanguage of the law. In this paper, we raise the issue and propose a trivial solution by introducing BERTLaw a legal sublanguage pretrained model. The paper's experiments demonstrate the superior effectiveness of the method compared to the baseline pretrained model

Keywords

Cite

@article{arxiv.2104.07782,
  title  = {Sublanguage: A Serious Issue Affects Pretrained Models in Legal Domain},
  author = {Ha-Thanh Nguyen and Le-Minh Nguyen},
  journal= {arXiv preprint arXiv:2104.07782},
  year   = {2021}
}