English

DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning

Artificial Intelligence 2026-04-22 v1 Databases

Abstract

This paper introduces DW-Bench, a new benchmark that evaluates large language models (LLMs) on graph-topology reasoning over data warehouse schemas, explicitly integrating both foreign-key (FK) and data-lineage edges. The benchmark comprises 1,046 automatically generated, verifiably correct questions across five schemas. Experiments show that tool-augmented methods substantially outperform static approaches but plateau on hard compositional subtypes.

Keywords

Cite

@article{arxiv.2604.18964,
  title  = {DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning},
  author = {Ahmed G. A. H Ahmed and C. Okan Sakar},
  journal= {arXiv preprint arXiv:2604.18964},
  year   = {2026}
}

Comments

24 pages, 6 figures. Datasets and evaluation code available at GitHub

R2 v1 2026-07-01T12:27:30.278Z