Generalised Bellman recurrence and three dualities in sequential decision-making
Abstract
What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through sufficient statistics, that the return decomposes recursively, and that the aggregation of uncertainty is compatible with both. When all three conditions hold on a common state, the Bellman equation arises from their mutual consistency; when one fails, tractability can often be recovered by augmenting the state or by deforming return or dynamics. The same conditions are shown to give rise to three dualities: one between probability and return, one between return and aggregation, and one between aggregation and probability. Our framework reveals these dualities as arising from a single construction, unifying methods developed separately across reinforcement learning, control, and decision theory.
Cite
@article{arxiv.2607.18077,
title = {Generalised Bellman recurrence and three dualities in sequential decision-making},
author = {Fernando E. Rosas and David Hyland and Daniel Polani},
journal= {arXiv preprint arXiv:2607.18077},
year = {2026}
}
Comments
27 pages, 2 figures