中文

熵梯度反转:向大规模推理模型内部机制迈进

人工智能 2026-05-25 v2 计算与语言

摘要

大规模推理模型(LRMs)的发展催化了从依赖“快速思维”文本生成向系统性、逐步推理的“慢思维”推理范式的转变,使其在复杂数学和逻辑任务中实现了最新水平的性能。然而,该领域面临着“token 级别行为分析与内部推理机制之间的根本鸿沟”以及“依赖高成本外部验证器的强化学习(RL)推理优化不稳定”这两大挑战。我们识别并正式定义了**熵梯度反转**(Entropy-Gradient Inversion),这是一种在标记熵和逻辑梯度之间保持稳健负相关的现象,作为 LRM 推理能力的明确几何指纹。基于此,我们提出了**关联正则化组策略优化(CorR-PO)**,将该反转签名嵌入 RL 奖励正则化中。广泛的实验表明,CorR-PO 在多个模型规模的各种推理基准上 consistently 优于最新基线,证明更强的反转直接与更优的推理性能相关。

关键词

引用

@article{arxiv.2605.17770,
  title  = {Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models},
  author = {Junyao Yang and Chen Qian and Kun Wang and Linfeng Zhang and Quanshi Zhang and Yong Liu and Dongrui Liu},
  journal= {arXiv preprint arXiv:2605.17770},
  year   = {2026}
}

备注

The authors are withdrawing this manuscript due to fundamental inaccuracies in the institutional affiliations and administrative attributions provided at the time of submission. As this version cannot be validated under the correct institutional framework, the authors request its formal withdrawal from the repository. No immediate replacement is intended