English

BSBench: will your LLM find the largest prime number?

Computation and Language 2025-06-06 v1

Abstract

We propose that benchmarking LLMs on questions which have no reasonable answer actually isn't as silly as it sounds. We also present a benchmark that allows such testing and a method to modify the existing datasets, and discover that existing models demonstrate a performance far from the perfect on such questions. Our code and data artifacts are available at https://github.com/L3G5/impossible-bench

Keywords

Cite

@article{arxiv.2506.04535,
  title  = {BSBench: will your LLM find the largest prime number?},
  author = {K. O. T. Erziev},
  journal= {arXiv preprint arXiv:2506.04535},
  year   = {2025}
}

Comments

7 + 2 pages

R2 v1 2026-07-01T03:00:22.130Z