English

What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models

Computation and Language 2025-07-31 v1 Artificial Intelligence

Abstract

Recent work has argued that large language models (LLMs) are not "abstract reasoners", citing their poor zero-shot performance on a variety of challenging tasks as evidence. We revisit these experiments in order to add nuance to the claim. First, we show that while LLMs indeed perform poorly in a zero-shot setting, even tuning a small subset of parameters for input encoding can enable near-perfect performance. However, we also show that this finetuning does not necessarily transfer across datasets. We take this collection of empirical results as an invitation to (re-)open the discussion of what it means to be an "abstract reasoner", and why it matters whether LLMs fit the bill.

Keywords

Cite

@article{arxiv.2507.22457,
  title  = {What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models},
  author = {Tian Yun and Chen Sun and Ellie Pavlick},
  journal= {arXiv preprint arXiv:2507.22457},
  year   = {2025}
}

Comments

CONLL 2025. Project webpage: https://abstract-reasoner-llm.github.io/

R2 v1 2026-07-01T04:25:30.812Z