English

Phase-Aware Mixture of Experts for Agentic Reinforcement Learning

Artificial Intelligence 2026-05-20 v3

Abstract

Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a \emph{single} policy network, causing \emph{simplicity bias} where simple tasks occupy most parameters and dominate gradient updates, leaving insufficient capacity for complex tasks. A plausible remedy could be employing the Mixture-of-Experts (MoE) architecture in the policy network, as MoE allows different parameters (experts) to specialize in different tasks, preventing simple tasks from dominating all parameters. However, a key limitation of traditional MoE is its token-level routing, where the router assigns each token to specialized experts, which fragments phase-consistent patterns into scattered expert assignments and thus undermines expert specialization. In this paper, we propose \textbf{Phase-Aware Mixture of Experts (PA-MoE)}. It first features a lightweight \emph{phase router} that learns latent phase boundaries directly from the RL objective without pre-defining phase categories. Then, the phase router allocates temporally consistent assignments to the same expert, allowing experts to preserve phase-specific expertise. Experimental results demonstrate the effectiveness of our proposed PA-MoE.

Keywords

Cite

@article{arxiv.2602.17038,
  title  = {Phase-Aware Mixture of Experts for Agentic Reinforcement Learning},
  author = {Shengtian Yang and Yu Li and Shuo He and Yewen Li and Qingpeng Cai and Peng Jiang and Lei Feng},
  journal= {arXiv preprint arXiv:2602.17038},
  year   = {2026}
}
R2 v1 2026-07-01T10:42:24.092Z