Breaking the Bias Barrier in Concave Multi-Objective Reinforcement Learning
Abstract
While standard reinforcement learning optimizes a single reward signal, many applications require optimizing a nonlinear utility over multiple objectives, where each denotes the expected discounted return of a distinct reward function. A common approach is concave scalarization, which captures important trade-offs such as fairness and risk sensitivity. However, nonlinear scalarization introduces a fundamental challenge for policy gradient methods: the gradient depends on , while in practice only empirical return estimates are available. Because is nonlinear, the plug-in estimator is biased (), leading to persistent gradient bias that degrades sample complexity. In this work we identify and overcome this bias barrier in concave-scalarized multi-objective reinforcement learning. We show that existing policy-gradient methods suffer an intrinsic sample complexity due to this bias. To address this issue, we develop a Natural Policy Gradient (NPG) algorithm equipped with a multi-level Monte Carlo (MLMC) estimator that controls the bias of the scalarization gradient while maintaining low sampling cost. We prove that this approach achieves the optimal sample complexity for computing an -optimal policy. Furthermore, we show that when the scalarization function is second-order smooth, the first-order bias cancels automatically, allowing vanilla NPG to achieve the same rate without MLMC. Our results provide the first optimal sample complexity guarantees for concave multi-objective reinforcement learning under policy-gradient methods.
Cite
@article{arxiv.2603.08518,
title = {Breaking the Bias Barrier in Concave Multi-Objective Reinforcement Learning},
author = {Swetha Ganesh and Vaneet Aggarwal},
journal= {arXiv preprint arXiv:2603.08518},
year = {2026}
}