English

Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning

Machine Learning 2025-02-07 v3

Abstract

An important challenge in multi-objective reinforcement learning is obtaining a Pareto front of policies to attain optimal performance under different preferences. We introduce Iterated Pareto Referent Optimisation (IPRO), which decomposes finding the Pareto front into a sequence of constrained single-objective problems. This enables us to guarantee convergence while providing an upper bound on the distance to undiscovered Pareto optimal solutions at each step. We evaluate IPRO using utility-based metrics and its hypervolume and find that it matches or outperforms methods that require additional assumptions. By leveraging problem-specific single-objective solvers, our approach also holds promise for applications beyond multi-objective reinforcement learning, such as planning and pathfinding.

Keywords

Cite

@article{arxiv.2402.07182,
  title  = {Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning},
  author = {Willem Röpke and Mathieu Reymond and Patrick Mannion and Diederik M. Roijers and Ann Nowé and Roxana Rădulescu},
  journal= {arXiv preprint arXiv:2402.07182},
  year   = {2025}
}

Comments

Accepted at AAMAS 2025

R2 v1 2026-06-28T14:45:18.531Z