English

PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning

Machine Learning 2025-09-24 v2

Abstract

Partial observability presents a significant challenge for Safe Reinforcement Learning (Safe RL), as it impedes the identification of potential risks and rewards. Leveraging specific types of privileged information during training to mitigate the effects of partial observability has yielded notable empirical successes. In this paper, we propose Asymmetric Constrained Partially Observable Markov Decision Processes (ACPOMDPs) to theoretically examine the advantages of incorporating privileged information in Safe RL. Building upon ACPOMDPs, we propose the Privileged Information Guided Dreamer (PIGDreamer), a model-based RL approach that leverages privileged information to enhance the agent's safety and performance through privileged representation alignment and an asymmetric actor-critic structure. Our empirical results demonstrate that PIGDreamer significantly outperforms existing Safe RL methods. Furthermore, compared to alternative privileged RL methods, our approach exhibits enhanced performance, robustness, and efficiency. Codes are available at: https://github.com/hggforget/PIGDreamer.

Keywords

Cite

@article{arxiv.2508.02159,
  title  = {PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning},
  author = {Dongchi Huang and Jiaqi Wang and Yang Li and Chunhe Xia and Tianle Zhang and Kaige Zhang},
  journal= {arXiv preprint arXiv:2508.02159},
  year   = {2025}
}

Comments

ICML 2025

R2 v1 2026-07-01T04:32:48.744Z