English

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

Artificial Intelligence 2026-07-13 v1

Abstract

Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria. We introduce SDABench, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics). SDABench comprises 527 real-data instances (SDA-Real) and 6000 synthetic instances (SDA-Synth), each in both multiple-choice and open-ended formats, constructed through an automated pipeline. Evaluating 15 representative LLMs, we find that models handle descriptive analysis well but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning. SDABench further provides a five-stage error analysis framework that locates where LLMs fail: more advanced models more reliably identify the relevant scope and variables, but still struggle to select appropriate analytical procedures, model variable relationships, and draw valid conclusions.

Cite

@article{arxiv.2607.11079,
  title  = {Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists},
  author = {Chuhan Shi and Xiaoquan Ren and Sicheng Song and Haobo Li and Rui Sheng and Yushi Sun},
  journal= {arXiv preprint arXiv:2607.11079},
  year   = {2026}
}
R2 v1 2026-07-22T20:36:41.385Z