English

Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization

Computation and Language 2026-04-14 v1 Artificial Intelligence Machine Learning

Abstract

To encourage diverse exploration in reinforcement learning (RL) for large language models (LLMs) without compromising accuracy, we propose Policy Split, a novel paradigm that bifurcates the policy into normal and high-entropy modes with a high-entropy prompt. While sharing model parameters, the two modes undergo collaborative dual-mode entropy regularization tailored to distinct objectives. Specifically, the normal mode optimizes for task correctness, while the high-entropy mode incorporates a preference for exploration, and the two modes learn collaboratively. Extensive experiments demonstrate that our approach consistently outperforms established entropy-guided RL baselines across various model sizes in general and creative tasks. Further analysis reveals that Policy Split facilitates dual-mode exploration, where the high-entropy mode generates distinct behavioral patterns to the normal mode, providing unique learning signals.

Keywords

Cite

@article{arxiv.2604.11510,
  title  = {Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization},
  author = {Jiashu Yao and Heyan Huang and Chuwei Luo and Daiqing Wu and Zeming Liu and Yuhang Guo and Yangyang Kang},
  journal= {arXiv preprint arXiv:2604.11510},
  year   = {2026}
}

Comments

preprint

R2 v1 2026-07-01T12:06:29.766Z