English

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

Computation and Language 2026-07-06 v1 Artificial Intelligence Computer Vision and Pattern Recognition

Abstract

Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications. However, this adoption is limited for low-resource languages such as Bangla due to the scarcity of high-quality annotated data. To address this gap, we introduce BaFCo, a benchmark dataset for Bangla form comprehension with a focus on Document Layout Analysis (DLA) and Key Information Extraction (KIE). BaFCo curates 200 multi-page complex Bangladeshi government forms, sourced from across diverse sectors including agriculture, education, banking, and land management. To accurately capture the structural and contextual complexity of these forms, we define a fine-grained annotation schema comprising 26 types of form entities, along with a separate coarse form entity set consisting of 5 types. We evaluate the latest MLLMs from the ChatGPT, Gemini, Claude, Qwen, and Kimi series using zero-shot and chain-of-thought prompts under both low and high reasoning setups. Our results reveal limitations in current MLLMs' ability in comprehending Bangla forms, particularly in accurately localizing highly granular form entities. Our dataset and code is available at: https://huggingface.co/datasets/Mausul/bafco

Cite

@article{arxiv.2607.05614,
  title  = {BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension},
  author = {Abu Tyeb Azad and Ishita Sur Apan and Fahim Ahmed and Sumaiya Karim Katha and Ezharuddin Jubaer and Armun Alam and Pranjal Kumar Nandi and Amin Ahsan Ali and Aman Chadha and Md Mofijul Islam and AKM Mahbubur Rahman},
  journal= {arXiv preprint arXiv:2607.05614},
  year   = {2026}
}

Comments

Accepted at the 19th European Conference on Computer Vision (ECCV), 2026