English

LeanCat: A Benchmark Suite for Formal Category Theory in Lean (Part I: 1-Categories)

Logic in Computer Science 2026-02-27 v2 Artificial Intelligence Formal Languages and Automata Theory Machine Learning Category Theory

Abstract

While large language models (LLMs) have demonstrated impressive capabilities in formal theorem proving, current benchmarks fail to adequately measure library-grounded abstraction -- the ability to reason with high-level interfaces and reusable structures central to modern mathematics and software engineering. We introduce LeanCat, a challenging benchmark comprising 100 fully formalized category-theory tasks in Lean. Unlike algebra or arithmetic, category theory serves as a rigorous stress test for structural, interface-level reasoning. Our evaluation reveals a severe abstraction gap: the best state-of-the-art model solves only 12.0% of tasks at pass@4, with performance collapsing from 55.0% on Easy tasks to 0.0% on High-difficulty tasks, highlighting a failure in compositional generalization. To overcome this, we evaluate LeanBridge, a retrieval-augmented agent that employs a retrieve-generate-verify loop. LeanBridge achieves a peak success rate of 24.0% -- doubling the performance of the best static baseline. These results empirically demonstrate that iterative refinement and dynamic library retrieval are not merely optimizations but strict necessities for neuro-symbolic reasoning in abstract domains. LeanCat offers a compact, reusable testbed for tracking progress toward reliable, research-level formalization.

Keywords

Cite

@article{arxiv.2512.24796,
  title  = {LeanCat: A Benchmark Suite for Formal Category Theory in Lean (Part I: 1-Categories)},
  author = {Rongge Xu and Hui Dai and Yiming Fu and Jiedong Jiang and Tianjiao Nie and Junkai Wang and Holiverse Yang and Zhi-Hao Zhang},
  journal= {arXiv preprint arXiv:2512.24796},
  year   = {2026}
}

Comments

22 pages, 9 figures, 5 tables