English

AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

Databases 2026-05-26 v1 Artificial Intelligence Machine Learning

Abstract

We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through \emph{latent world recovery}. AvalancheBench improves on existing benchmarks in three ways. First, it evaluates analytical understanding rather than pipeline completion: systems are scored on whether they recover the segments, drivers, temporal events, and relationships that explain the data, not merely on whether they execute a workflow or produce a plausible report. Second, it provides ground truth for goal-driven analytics by generating observations from a known latent world, enabling partial credit for incomplete but valid recoveries. Third, it exposes how early analytical mistakes propagate into later conclusions: missed segments, merged events, or wrong attributions can lead to systematically wrong recommendations. In this sense, AvalancheBench complements real-data benchmarks by providing a controlled setting for diagnosing whether agents recover the analytical structure behind enterprise data. On a first e-commerce use case, the strongest configuration of a leading coding agent recovers only 26\% of the rubric, with failures concentrated in generic customer segmentations and merged temporal events.

Keywords

Cite

@article{arxiv.2605.24183,
  title  = {AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery},
  author = {Darek Kleczek and Fuheng Zhao and Alexander W. Lee and Julien Tissier and Pawel Liskowski and Ugur Cetintemel and Anupam Datta},
  journal= {arXiv preprint arXiv:2605.24183},
  year   = {2026}
}
R2 v1 2026-07-22T07:29:24.278Z