English

Arctic-Extract Technical Report

Computation and Language 2025-11-21 v1 Computer Vision and Pattern Recognition

Abstract

Arctic-Extract is a state-of-the-art model designed for extracting structural data (question answering, entities and tables) from scanned or digital-born business documents. Despite its SoTA capabilities, the model is deployable on resource-constrained hardware, weighting only 6.6 GiB, making it suitable for deployment on devices with limited resources, such as A10 GPUs with 24 GB of memory. Arctic-Extract can process up to 125 A4 pages on those GPUs, making suitable for long document processing. This paper highlights Arctic-Extract's training protocols and evaluation results, demonstrating its strong performance in document understanding.

Keywords

Cite

@article{arxiv.2511.16470,
  title  = {Arctic-Extract Technical Report},
  author = {Mateusz Chiliński and Julita Ołtusek and Wojciech Jaśkowski},
  journal= {arXiv preprint arXiv:2511.16470},
  year   = {2025}
}
R2 v1 2026-07-01T07:47:28.346Z