English

Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature

Computation and Language 2024-12-18 v1 Computer Vision and Pattern Recognition Machine Learning

Abstract

Scientific Literature charts often contain complex visual elements, including multi-plot figures, flowcharts, structural diagrams and etc. Evaluating multimodal models using these authentic and intricate charts provides a more accurate assessment of their understanding abilities. However, existing benchmarks face limitations: a narrow range of chart types, overly simplistic template-based questions and visual elements, and inadequate evaluation methods. These shortcomings lead to inflated performance scores that fail to hold up when models encounter real-world scientific charts. To address these challenges, we introduce a new benchmark, Scientific Chart QA (SCI-CQA), which emphasizes flowcharts as a critical yet often overlooked category. To overcome the limitations of chart variety and simplistic visual elements, we curated a dataset of 202,760 image-text pairs from 15 top-tier computer science conferences papers over the past decade. After rigorous filtering, we refined this to 37,607 high-quality charts with contextual information. SCI-CQA also introduces a novel evaluation framework inspired by human exams, encompassing 5,629 carefully curated questions, both objective and open-ended. Additionally, we propose an efficient annotation pipeline that significantly reduces data annotation costs. Finally, we explore context-based chart understanding, highlighting the crucial role of contextual information in solving previously unanswerable questions.

Keywords

Cite

@article{arxiv.2412.12150,
  title  = {Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature},
  author = {Lingdong Shen and Qigqi and Kun Ding and Gaofeng Meng and Shiming Xiang},
  journal= {arXiv preprint arXiv:2412.12150},
  year   = {2024}
}
R2 v1 2026-06-28T20:37:39.164Z