English

Watson & Holmes: A Naturalistic Benchmark for Comparing Human and LLM Reasoning

Artificial Intelligence 2026-02-24 v1

Abstract

Existing benchmarks for AI reasoning provide limited insight into how closely these capabilities resemble human reasoning in naturalistic contexts. We present an adaptation of the Watson & Holmes detective tabletop game as a new benchmark designed to evaluate reasoning performance using incrementally presented narrative evidence, open-ended questions and unconstrained language responses. An automated grading system was developed and validated against human assessors to enable scalable and replicable performance evaluation. Results show a clear improvement in AI model performance over time. Over nine months of 2025, model performance rose from the lower quartile of the human comparison group to approximately the top 5%. Around half of this improvement reflects steady advancement across successive model releases, while the remainder corresponds to a marked step change associated with reasoning-oriented model architectures. Systematic differences in the performance of AI models compared to humans, dependent on features of the specific detection puzzle, were mostly absent with the exception of a fall in performance for models when solving longer cases (case lengths being in the range of 1900-4000 words), and an advantage at inductive reasoning for reasoning models at early stages of case solving when evidence was scant.

Keywords

Cite

@article{arxiv.2602.19914,
  title  = {Watson & Holmes: A Naturalistic Benchmark for Comparing Human and LLM Reasoning},
  author = {Thatchawin Leelawat and Lewis D Griffin},
  journal= {arXiv preprint arXiv:2602.19914},
  year   = {2026}
}

Comments

51 pages, 13 figures

R2 v1 2026-07-01T10:47:31.536Z