English

LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

Artificial Intelligence 2025-07-16 v3 Computation and Language

Abstract

Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wider range of tasks. However, existing document understanding benchmarks have been limited to handling only a small number of pages and fail to provide a comprehensive analysis of layout elements locating. In this paper, we first define three primary task categories: Long Document Understanding, numerical Reasoning, and cross-element Locating, and then propose a comprehensive benchmark, LongDocURL, integrating above three primary tasks and comprising 20 sub-tasks categorized based on different primary tasks and answer evidences. Furthermore, we develop a semi-automated construction pipeline and collect 2,325 high-quality question-answering pairs, covering more than 33,000 pages of documents, significantly outperforming existing benchmarks. Subsequently, we conduct comprehensive evaluation experiments on both open-source and closed-source models across 26 different configurations, revealing critical performance gaps in this field.

Keywords

Cite

@article{arxiv.2412.18424,
  title  = {LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating},
  author = {Chao Deng and Jiale Yuan and Pi Bu and Peijie Wang and Zhong-Zhi Li and Jian Xu and Xiao-Hui Li and Yuan Gao and Jun Song and Bo Zheng and Cheng-Lin Liu},
  journal= {arXiv preprint arXiv:2412.18424},
  year   = {2025}
}
R2 v1 2026-06-28T20:48:04.571Z