English

A Survey of Reinforcement Learning for Large Reasoning Models

Computation and Language 2025-10-10 v3 Artificial Intelligence Machine Learning

Abstract

In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs). RL has achieved remarkable success in advancing the frontier of LLM capabilities, particularly in addressing complex logical tasks such as mathematics and coding. As a result, RL has emerged as a foundational methodology for transforming LLMs into LRMs. With the rapid progress of the field, further scaling of RL for LRMs now faces foundational challenges not only in computational resources but also in algorithm design, training data, and infrastructure. To this end, it is timely to revisit the development of this domain, reassess its trajectory, and explore strategies to enhance the scalability of RL toward Artificial SuperIntelligence (ASI). In particular, we examine research applying RL to LLMs and LRMs for reasoning abilities, especially since the release of DeepSeek-R1, including foundational components, core problems, training resources, and downstream applications, to identify future opportunities and directions for this rapidly evolving area. We hope this review will promote future research on RL for broader reasoning models. Github: https://github.com/TsinghuaC3I/Awesome-RL-for-LRMs

Keywords

Cite

@article{arxiv.2509.08827,
  title  = {A Survey of Reinforcement Learning for Large Reasoning Models},
  author = {Kaiyan Zhang and Yuxin Zuo and Bingxiang He and Youbang Sun and Runze Liu and Che Jiang and Yuchen Fan and Kai Tian and Guoli Jia and Pengfei Li and Yu Fu and Xingtai Lv and Yuchen Zhang and Sihang Zeng and Shang Qu and Haozhan Li and Shijie Wang and Yuru Wang and Xinwei Long and Fangfu Liu and Xiang Xu and Jiaze Ma and Xuekai Zhu and Ermo Hua and Yihao Liu and Zonglin Li and Huayu Chen and Xiaoye Qu and Yafu Li and Weize Chen and Zhenzhao Yuan and Junqi Gao and Dong Li and Zhiyuan Ma and Ganqu Cui and Zhiyuan Liu and Biqing Qi and Ning Ding and Bowen Zhou},
  journal= {arXiv preprint arXiv:2509.08827},
  year   = {2025}
}

Comments

Fixed typos; added missing and recent citations (117 -> 120 pages)