English

Goal-Conditioned Reinforcement Learning from Sub-Optimal Data on Metric Spaces

Machine Learning 2026-02-12 v3

Abstract

We study the problem of learning optimal behavior from sub-optimal datasets for goal-conditioned offline reinforcement learning under sparse rewards, invertible actions and deterministic transitions. To mitigate the effects of \emph{distribution shift}, we propose MetricRL, a method that combines metric learning for value function approximation with weighted imitation learning for policy estimation. MetricRL avoids conservative or behavior-cloning constraints, enabling effective learning even in severely sub-optimal regimes. We introduce distance monotonicity as a key property linking metric representations to optimality and design an objective that explicitly promotes it. Empirically, MetricRL consistently outperforms prior state-of-the-art goal-conditioned RL methods in recovering near-optimal behavior from sub-optimal offline data.

Keywords

Cite

@article{arxiv.2402.10820,
  title  = {Goal-Conditioned Reinforcement Learning from Sub-Optimal Data on Metric Spaces},
  author = {Alfredo Reichlin and Miguel Vasco and Hang Yin and Danica Kragic},
  journal= {arXiv preprint arXiv:2402.10820},
  year   = {2026}
}
R2 v1 2026-06-28T14:50:54.538Z