English

VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking

Human-Computer Interaction 2026-03-02 v2

Abstract

Recent advances have showcased the extraordinary capabilities of Large Language Model (LLM) agents in tackling web-based information-seeking tasks. However, existing efforts mainly focus on single-fact retrieval and rely on outcome-only verification, thereby limiting their scalability in realistic knowledge-intensive scenarios that involve long-horizon web tasks requiring large-scale retrieval and synthesis of information from diverse sources. In this work, we introduce VeriWeb, a novel verifiable long-chain web benchmark designed to facilitate the evaluation and development of web agents within realistic web environments. Our benchmark emphasizes two critical dimensions: (1) long-chain complexity, encompassing both breadth- and depth-oriented search tasks to assess how effectively web agents ensure comprehensive information coverage and consistent context tracking in multi-hop reasoning; and (2) subtask-level verifiability, where tasks are decomposed into a sequence of interdependent verifiable subtasks. This structure enables diverse exploration strategies within each subtask, while ensuring that each subtask-level answer remains unchanged and verifiable. The benchmark consists of 302 tasks across five real-world domains, each with a complete trajectory demonstration, annotated by human experts. Extensive experiments on VeriWeb using various agents powered by different foundation models reveal significant performance gaps in handling long-horizon web tasks, highlighting the need for more powerful agentic information-seeking capabilities.

Keywords

Cite

@article{arxiv.2508.04026,
  title  = {VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking},
  author = {Shunyu Liu and Minghao Liu and Huichi Zhou and Zhenyu Cui and Yang Zhou and Yuhao Zhou and Jialiang Gao and Heng Zhou and Yunhao Yang and Wendong Fan and puzhen zhang and Ge Zhang and Jiajun Shi and Weihao Xuan and Jiaxing Huang and Shuang Luo and Fang Wu and Heli Qi and Qingcheng Zeng and Junjie Wang and Aosong Feng and Jindi Lv and Sicong Jiang and Ziqi Ren and Wangchunshu Zhou and Zhenfei Yin and Wenlong Zhang and Guohao Li and Wenhao Yu and Lei Ma and Lei Bai and Qunshu Lin and Mingli Song and Dacheng Tao},
  journal= {arXiv preprint arXiv:2508.04026},
  year   = {2026}
}
R2 v1 2026-07-01T04:36:26.989Z