English

Training LayoutLM from Scratch for Efficient Named-Entity Recognition in the Insurance Domain

Computation and Language 2024-12-13 v1

Abstract

Generic pre-trained neural networks may struggle to produce good results in specialized domains like finance and insurance. This is due to a domain mismatch between training data and downstream tasks, as in-domain data are often scarce due to privacy constraints. In this work, we compare different pre-training strategies for LayoutLM. We show that using domain-relevant documents improves results on a named-entity recognition (NER) problem using a novel dataset of anonymized insurance-related financial documents called Payslips. Moreover, we show that we can achieve competitive results using a smaller and faster model.

Keywords

Cite

@article{arxiv.2412.09341,
  title  = {Training LayoutLM from Scratch for Efficient Named-Entity Recognition in the Insurance Domain},
  author = {Benno Uthayasooriyar and Antoine Ly and Franck Vermet and Caio Corro},
  journal= {arXiv preprint arXiv:2412.09341},
  year   = {2024}
}

Comments

Coling 2025 workshop (FinNLP)

R2 v1 2026-06-28T20:32:35.304Z