English

SFBench: The SciFy Scientific Feasibility Benchmark

Artificial Intelligence 2026-06-28 v1

Abstract

We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claims in materials science, each annotated with a ground-truth feasibility score on a five-point scale along with an explanation of that assessment. The collection differs from previous collections in several important ways: 1) it defines a complex task that requires reasoning over claims of varying scientific feasibility; 2) its claims are not extracted from existing scientific publications but are created de novo, greatly reducing the chances that LLMs have trained on them; 3) claims and ground truth are established by subject matter experts, not by artificial intelligence; and 4) unlike many benchmarks that ask about question/answer pairs, provide multiple choice answers, or ask questions requiring short, fixed answers, SFBench explanations are completely open-ended. We describe the benchmark design, data creation process, and evaluation metrics, and we report baseline results using recent GPT models.

Cite

@article{arxiv.2606.29630,
  title  = {SFBench: The SciFy Scientific Feasibility Benchmark},
  author = {Cash Costello and James Mayfield and Elsbeth Turcan and Christine Piatko and Christina K. Pikas and Justin Rokisky and Sam Scheck and Chris Ribaudo and Ritwik Bose and Alex Memory},
  journal= {arXiv preprint arXiv:2606.29630},
  year   = {2026}
}