English

Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza Pendulum

Machine Learning 2026-07-15 v1

Abstract

We suggest using the Lyapunov characteristic exponent (LCE) as a dense reward signal for the reinforcement learning problem of stabilizing the inverted pendulum with vertical motion. With LCE, the agent not only successfully found the oscillatory motion known as the Kapitza pendulum but also damped the pendulum's pivoting, leaving it in a strictly upright position.

Keywords

Cite

@article{arxiv.2607.14001,
  title  = {Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza Pendulum},
  author = {Slava Andrejev},
  journal= {arXiv preprint arXiv:2607.14001},
  year   = {2026}
}