English

Golden Handcuffs make safer AI agents

Machine Learning 2026-04-16 v1 Artificial Intelligence

Abstract

Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand the agent's subjective reward range to include a large negative value L-L, while the true environment's rewards lie in [0,1][0,1]. After observing consistently high rewards, the Bayesian policy becomes risk-averse to novel schemes that plausibly lead to L-L. We design a simple override mechanism that yields control to a safe mentor whenever the predicted value drops below a fixed threshold. We prove two properties of the resulting agent: (i) Capability: using mentor-guided exploration with vanishing frequency, the agent attains sublinear regret against its best mentor. (ii) Safety: no decidable low-complexity predicate is triggered by the optimizing policy before it is triggered by a mentor.

Keywords

Cite

@article{arxiv.2604.13609,
  title  = {Golden Handcuffs make safer AI agents},
  author = {Aram Ebtekar and Michael K. Cohen},
  journal= {arXiv preprint arXiv:2604.13609},
  year   = {2026}
}

Comments

26 pages, preliminary version

R2 v1 2026-07-01T12:10:20.342Z