English

Robust Exploration in Directed Controller Synthesis via Reinforcement Learning with Soft Mixture-of-Experts

Artificial Intelligence 2026-02-24 v1 Machine Learning

Abstract

On-the-fly Directed Controller Synthesis (OTF-DCS) mitigates state-space explosion by incrementally exploring the system and relies critically on an exploration policy to guide search efficiently. Recent reinforcement learning (RL) approaches learn such policies and achieve promising zero-shot generalization from small training instances to larger unseen ones. However, a fundamental limitation is anisotropic generalization, where an RL policy exhibits strong performance only in a specific region of the domain-parameter space while remaining fragile elsewhere due to training stochasticity and trajectory-dependent bias. To address this, we propose a Soft Mixture-of-Experts framework that combines multiple RL experts via a prior-confidence gating mechanism and treats these anisotropic behaviors as complementary specializations. The evaluation on the Air Traffic benchmark shows that Soft-MoE substantially expands the solvable parameter space and improves robustness compared to any single expert.

Keywords

Cite

@article{arxiv.2602.19244,
  title  = {Robust Exploration in Directed Controller Synthesis via Reinforcement Learning with Soft Mixture-of-Experts},
  author = {Toshihide Ubukata and Zhiyao Wang and Enhong Mu and Jialong Li and Kenji Tei},
  journal= {arXiv preprint arXiv:2602.19244},
  year   = {2026}
}
R2 v1 2026-07-01T10:46:24.298Z