On Path to Multimodal Historical Reasoning: HistBench and HistAgent
Abstract
Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored. Historical reasoning poses unique challenges for AI, involving multimodal source interpretation, temporal inference, and cross-linguistic analysis. While general-purpose agents perform well on many existing benchmarks, they lack the domain-specific expertise required to engage with historical materials and questions. To address this gap, we introduce HistBench, a new benchmark of 414 high-quality questions designed to evaluate AI's capacity for historical reasoning and authored by more than 40 expert contributors. The tasks span a wide range of historical problems-from factual retrieval based on primary sources to interpretive analysis of manuscripts and images, to interdisciplinary challenges involving archaeology, linguistics, or cultural history. Furthermore, the benchmark dataset spans 29 ancient and modern languages and covers a wide range of historical periods and world regions. Finding the poor performance of LLMs and other agents on HistBench, we further present HistAgent, a history-specific agent equipped with carefully designed tools for OCR, translation, archival search, and image understanding in History. On HistBench, HistAgent based on GPT-4o achieves an accuracy of 27.54% pass@1 and 36.47% pass@2, significantly outperforming LLMs with online search and generalist agents, including GPT-4o (18.60%), DeepSeek-R1(14.49%) and Open Deep Research-smolagents(20.29% pass@1 and 25.12% pass@2). These results highlight the limitations of existing LLMs and generalist agents and demonstrate the advantages of HistAgent for historical reasoning.
Keywords
Cite
@article{arxiv.2505.20246,
title = {On Path to Multimodal Historical Reasoning: HistBench and HistAgent},
author = {Jiahao Qiu and Fulian Xiao and Yimin Wang and Yuchen Mao and Yijia Chen and Xinzhe Juan and Shu Zhang and Siran Wang and Xuan Qi and Tongcheng Zhang and Zixin Yao and Jiacheng Guo and Yifu Lu and Charles Argon and Jundi Cui and Daixin Chen and Junran Zhou and Shuyao Zhou and Zhanpeng Zhou and Ling Yang and Shilong Liu and Hongru Wang and Kaixuan Huang and Xun Jiang and Yuming Cao and Yue Chen and Yunfei Chen and Zhengyi Chen and Ruowei Dai and Mengqiu Deng and Jiye Fu and Yunting Gu and Zijie Guan and Zirui Huang and Xiaoyan Ji and Yumeng Jiang and Delong Kong and Haolong Li and Jiaqi Li and Ruipeng Li and Tianze Li and Zhuoran Li and Haixia Lian and Mengyue Lin and Xudong Liu and Jiayi Lu and Jinghan Lu and Wanyu Luo and Ziyue Luo and Zihao Pu and Zhi Qiao and Ruihuan Ren and Liang Wan and Ruixiang Wang and Tianhui Wang and Yang Wang and Zeyu Wang and Zihua Wang and Yujia Wu and Zhaoyi Wu and Hao Xin and Weiao Xing and Ruojun Xiong and Weijie Xu and Yao Shu and Yao Xiao and Xiaorui Yang and Yuchen Yang and Nan Yi and Jiadong Yu and Yangyuxuan Yu and Huiting Zeng and Danni Zhang and Yunjie Zhang and Zhaoyu Zhang and Zhiheng Zhang and Xiaofeng Zheng and Peirong Zhou and Linyan Zhong and Xiaoyin Zong and Ying Zhao and Zhenxin Chen and Lin Ding and Xiaoyu Gao and Bingbing Gong and Yichao Li and Yang Liao and Guang Ma and Tianyuan Ma and Xinrui Sun and Tianyi Wang and Han Xia and Ruobing Xian and Gen Ye and Tengfei Yu and Wentao Zhang and Yuxi Wang and Xi Gao and Mengdi Wang},
journal= {arXiv preprint arXiv:2505.20246},
year = {2025}
}
Comments
17 pages, 7 figures