English

Arctic-TILT. Business Document Understanding at Sub-Billion Scale

Computation and Language 2024-08-09 v1 Computer Vision and Pattern Recognition

Abstract

The vast portion of workloads employing LLMs involves answering questions grounded on PDF or scan content. We introduce the Arctic-TILT achieving accuracy on par with models 1000×\times its size on these use cases. It can be fine-tuned and deployed on a single 24GB GPU, lowering operational costs while processing Visually Rich Documents with up to 400k tokens. The model establishes state-of-the-art results on seven diverse Document Understanding benchmarks, as well as provides reliable confidence scores and quick inference, which are essential for processing files in large-scale or time-sensitive enterprise environments.

Keywords

Cite

@article{arxiv.2408.04632,
  title  = {Arctic-TILT. Business Document Understanding at Sub-Billion Scale},
  author = {Łukasz Borchmann and Michał Pietruszka and Wojciech Jaśkowski and Dawid Jurkiewicz and Piotr Halama and Paweł Józiak and Łukasz Garncarek and Paweł Liskowski and Karolina Szyndler and Andrzej Gretkowski and Julita Ołtusek and Gabriela Nowakowska and Artur Zawłocki and Łukasz Duhr and Paweł Dyda and Michał Turski},
  journal= {arXiv preprint arXiv:2408.04632},
  year   = {2024}
}
R2 v1 2026-06-28T18:07:58.795Z