English

Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models

Computation and Language 2026-01-21 v1

Abstract

Long Chain-of-Thought (LCoT), achieved by Reinforcement Learning with Verifiable Rewards (RLVR), has proven effective in enhancing the reasoning capabilities of Large Language Models (LLMs). However, reasoning in current LLMs is primarily generated as plain text, where performing semantic evaluation on such unstructured data creates a computational bottleneck during training. Despite RLVR-based optimization, existing methods still suffer from coarse-grained supervision, reward hacking, high training costs, and poor generalization. To address these issues, we propose the Graph Reasoning Paradigm (GRP), which realizes structured and symbolic reasoning, implemented via graph-structured representations with step-level cognitive labels. Building upon GRP, we further design Process-Aware Stratified Clipping Group Relative Policy Optimization (PASC-GRPO), which leverages structured evaluation to replace semantic evaluation, achieves process-aware verification through graph-structured outcome rewards, and mitigates reward hacking via stratified clipping advantage estimation. Experiments demonstrate significant improvements across mathematical reasoning and code generation tasks. Data, models, and code will be released later.

Keywords

Cite

@article{arxiv.2601.12995,
  title  = {Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models},
  author = {Runxuan Liu and Xianhao Ou and Xinyan Ma and Jiyuan Wang and Jiafeng Liang and Jiaqi Li and Tao He and Zheng Chu and Rongchuan Mu and Zekun Wang and Baoxin Wang and Dayong Wu and Ming Liu and Shijin Wang and Guoping Hu and Bing Qin},
  journal= {arXiv preprint arXiv:2601.12995},
  year   = {2026}
}