English

Why Can Large Language Models Generate Correct Chain-of-Thoughts?

Computation and Language 2024-06-07 v4

Abstract

This paper delves into the capabilities of large language models (LLMs), specifically focusing on advancing the theoretical comprehension of chain-of-thought prompting. We investigate how LLMs can be effectively induced to generate a coherent chain of thoughts. To achieve this, we introduce a two-level hierarchical graphical model tailored for natural language generation. Within this framework, we establish a compelling geometrical convergence rate that gauges the likelihood of an LLM-generated chain of thoughts compared to those originating from the true language. Our findings provide a theoretical justification for the ability of LLMs to produce the correct sequence of thoughts (potentially) explaining performance gains in tasks demanding reasoning skills.

Keywords

Cite

@article{arxiv.2310.13571,
  title  = {Why Can Large Language Models Generate Correct Chain-of-Thoughts?},
  author = {Rasul Tutunov and Antoine Grosnit and Juliusz Ziomek and Jun Wang and Haitham Bou-Ammar},
  journal= {arXiv preprint arXiv:2310.13571},
  year   = {2024}
}
R2 v1 2026-06-28T12:56:57.936Z