English

Towards Logically Sound Natural Language Reasoning with Logic-Enhanced Language Model Agents

Artificial Intelligence 2025-05-30 v2 Computation and Language Computer Science and Game Theory Logic in Computer Science

Abstract

Large language models (LLMs) are increasingly explored as general-purpose reasoners, particularly in agentic contexts. However, their outputs remain prone to mathematical and logical errors. This is especially challenging in open-ended tasks, where unstructured outputs lack explicit ground truth and may contain subtle inconsistencies. To address this issue, we propose Logic-Enhanced Language Model Agents (LELMA), a framework that integrates LLMs with formal logic to enable validation and refinement of natural language reasoning. LELMA comprises three components: an LLM-Reasoner, an LLM-Translator, and a Solver, and employs autoformalization to translate reasoning into logic representations, which are then used to assess logical validity. Using game-theoretic scenarios such as the Prisoner's Dilemma as testbeds, we highlight the limitations of both less capable (Gemini 1.0 Pro) and advanced (GPT-4o) models in generating logically sound reasoning. LELMA achieves high accuracy in error detection and improves reasoning correctness via self-refinement, particularly in GPT-4o. The study also highlights challenges in autoformalization accuracy and in evaluation of inherently ambiguous open-ended reasoning tasks.

Keywords

Cite

@article{arxiv.2408.16081,
  title  = {Towards Logically Sound Natural Language Reasoning with Logic-Enhanced Language Model Agents},
  author = {Agnieszka Mensfelt and Kostas Stathis and Vince Trencsenyi},
  journal= {arXiv preprint arXiv:2408.16081},
  year   = {2025}
}

Comments

Source code: https://github.com/dicelab-rhul/LELMA

R2 v1 2026-06-28T18:27:00.660Z