English

MUR: Momentum Uncertainty guided Reasoning for Large Language Models

Computation and Language 2026-05-26 v5

Abstract

Large Language Models have achieved impressive performance on reasoning-intensive tasks, yet optimizing their reasoning efficiency remains an open challenge. While Test-Time Scaling (TTS) improves reasoning quality, it often leads to overthinking, wasting tokens on redundant computations. This work investigates how to efficiently and adaptively guide current model' test-time scaling without additional training. Inspired by the concept of momentum in physics, we propose Momentum Uncertainty-guided Reasoning (MUR), which dynamically allocates thinking budgets to critical reasoning steps by tracking and aggregating stepwise uncertainty over time. To support flexible inference-time control, we introduce gamma-control, a simple mechanism that tunes the reasoning budget via a single hyperparameter. We provide in-depth theoretical proof to support the superiority of MUR in terms of stability and biases. MUR is comprehensively evaluated against various TTS methods across four challenging benchmarks (MATH-500, AIME24, AIME25, and GPQA-diamond) using different sizes of recent Qwen3 models (1.7B, 4B, and 8B). Results demonstrate that MUR reduces computation by by over 45% on average while improving accuracy from 0.33 to 3.46%.

Keywords

Cite

@article{arxiv.2507.14958,
  title  = {MUR: Momentum Uncertainty guided Reasoning for Large Language Models},
  author = {Hang Yan and Fangzhi Xu and Rongman Xu and Yifei Li and Jian Zhang and Haoran Luo and Xiaobao Wu and Luu Anh Tuan and Haiteng Zhao and Qika Lin and Jun Liu},
  journal= {arXiv preprint arXiv:2507.14958},
  year   = {2026}
}
R2 v1 2026-07-01T04:09:56.797Z