English

ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents

Artificial Intelligence 2024-09-24 v1 Computation and Language Computer Vision and Pattern Recognition Information Retrieval

Abstract

Multimodal key information extraction (KIE) models have been studied extensively on semi-structured documents. However, their investigation on unstructured documents is an emerging research topic. The paper presents an approach to adapt a multimodal transformer (i.e., ViBERTgrid previously explored on semi-structured documents) for unstructured financial documents, by incorporating a BiLSTM-CRF layer. The proposed ViBERTgrid BiLSTM-CRF model demonstrates a significant improvement in performance (up to 2 percentage points) on named entity recognition from unstructured documents in financial domain, while maintaining its KIE performance on semi-structured documents. As an additional contribution, we publicly released token-level annotations for the SROIE dataset in order to pave the way for its use in multimodal sequence labeling models.

Keywords

Cite

@article{arxiv.2409.15004,
  title  = {ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents},
  author = {Furkan Pala and Mehmet Yasin Akpınar and Onur Deniz and Gülşen Eryiğit},
  journal= {arXiv preprint arXiv:2409.15004},
  year   = {2024}
}

Comments

Accepted in MIDAS (The 8th Workshop on MIning DAta for financial applicationS) workshop of ECML PKDD 2023 conference

R2 v1 2026-06-28T18:53:41.999Z