English

Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings

Computation and Language 2024-07-31 v1

Abstract

We present Knesset-DictaBERT, a large Hebrew language model fine-tuned on the Knesset Corpus, which comprises Israeli parliamentary proceedings. The model is based on the DictaBERT architecture and demonstrates significant improvements in understanding parliamentary language according to the MLM task. We provide a detailed evaluation of the model's performance, showing improvements in perplexity and accuracy over the baseline DictaBERT model.

Keywords

Cite

@article{arxiv.2407.20581,
  title  = {Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings},
  author = {Gili Goldin and Shuly Wintner},
  journal= {arXiv preprint arXiv:2407.20581},
  year   = {2024}
}

Comments

3 pages, 1 table