English

Data-Driven Synthesis of Probabilistic Controlled Invariant Sets for Linear MDPs

Systems and Control 2026-04-06 v1 Systems and Control

Abstract

We study data-driven computation of probabilistic controlled invariant sets (PCIS) for safety-critical reinforcement learning under unknown dynamics. Assuming a linear MDP model, we use regularized least squares and self-normalized confidence bounds to construct a conservative estimate of the states from which the system can be kept inside a prescribed safe region over an NN-step horizon, together with the corresponding set-valued safe action map. This construction is obtained through a backward recursion and can be interpreted as a conservative approximation of the NN-step safety predecessor operator. When the associated conservative-inclusion event holds, a conservative fixed point of the approximate recursion can be certified as an (N,ϵ)(N,\epsilon)-PCIS with confidence at least η\eta. For continuous state spaces, we introduce a lattice abstraction and a Lipschitz-based discretization error bound to obtain a tractable approximation scheme. Finally, we use the resulting conservative fixed-point approximation as a runtime candidate PCIS in a practical shielding architecture with iterative updates, and illustrate the approach on a numerical experiment.

Keywords

Cite

@article{arxiv.2604.02727,
  title  = {Data-Driven Synthesis of Probabilistic Controlled Invariant Sets for Linear MDPs},
  author = {Kazumune Hashimoto and Shunki Kimura and Kazunobu Serizawa and Junya Ikemoto and Yulong Gao and Kai Cai},
  journal= {arXiv preprint arXiv:2604.02727},
  year   = {2026}
}
R2 v1 2026-07-01T11:52:21.104Z