English

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach

Machine Learning 2025-07-30 v1

Abstract

In many multi-objective reinforcement learning (MORL) applications, being able to systematically explore the Pareto-stationary solutions under multiple non-convex reward objectives with theoretical finite-time sample complexity guarantee is an important and yet under-explored problem. This motivates us to take the first step and fill the important gap in MORL. Specifically, in this paper, we propose a \uline{M}ulti-\uline{O}bjective weighted-\uline{CH}ebyshev \uline{A}ctor-critic (MOCHA) algorithm for MORL, which judiciously integrates the weighted-Chebychev (WC) and actor-critic framework to enable Pareto-stationarity exploration systematically with finite-time sample complexity guarantee. Sample complexity result of MOCHA algorithm reveals an interesting dependency on pminp_{\min} in finding an ϵ\epsilon-Pareto-stationary solution, where pminp_{\min} denotes the minimum entry of a given weight vector p\mathbf{p} in WC-scarlarization. By carefully choosing learning rates, the sample complexity for each exploration can be O~(ϵ2)\tilde{\mathcal{O}}(\epsilon^{-2}). Furthermore, simulation studies on a large KuaiRand offline dataset, show that the performance of MOCHA algorithm significantly outperforms other baseline MORL approaches.

Keywords

Cite

@article{arxiv.2507.21397,
  title  = {Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach},
  author = {Fnu Hairi and Jiao Yang and Tianchen Zhou and Haibo Yang and Chaosheng Dong and Fan Yang and Michinari Momma and Yan Gao and Jia Liu},
  journal= {arXiv preprint arXiv:2507.21397},
  year   = {2025}
}