English

Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding

Artificial Intelligence 2025-09-27 v4 Human-Computer Interaction

Abstract

Guesstimation -- the task of making approximate quantitative estimates about objects or events -- is a common real-world skill, yet remains underexplored in large language model (LLM) research. We introduce three guesstimation datasets: MARBLES, FUTURE, and ELECPRED, spanning physical estimation (e.g., how many marbles fit in a cup) to abstract predictions (e.g., the 2024 U.S. presidential election). Inspired by the social science concept of Wisdom of Crowds (WOC)- where the median of multiple estimates improves accuracy-we propose WOC decoding for LLMs. We replicate WOC effects in human participants and find that LLMs exhibit similar benefits: median aggregation across sampled responses consistently improves accuracy over greedy decoding, self-consistency decoding, and mean decoding. This suggests that LLMs encode a world model that supports approximate reasoning. Our results position guesstimation as a useful probe of LLM world knowledge and highlight WOC decoding as a strategy for enhancing LLM guesstimation performance on real-world tasks.

Keywords

Cite

@article{arxiv.2501.17310,
  title  = {Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding},
  author = {Yun-Shiuan Chuang and Sameer Narendran and Nikunj Harlalka and Alexander Cheung and Sizhe Gao and Siddharth Suresh and Junjie Hu and Timothy T. Rogers},
  journal= {arXiv preprint arXiv:2501.17310},
  year   = {2025}
}
R2 v1 2026-06-28T21:22:57.853Z