English

Back to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning

Machine Learning 2026-05-07 v2

Abstract

A recurring challenge in preference fine-tuning (PFT) is handling intransitive\textit{intransitive} (i.e., cyclic) preferences. Intransitive preferences often stem from either (i)\textit{(i)} inconsistent rankings along a single objective or (ii)\textit{(ii)} scalarizing multiple objectives into a single metric. Regardless of their source, the downstream implication of intransitive preferences is the same: there is no well-defined optimal policy, breaking a core assumption of the standard PFT pipeline. In response, we propose a novel, game-theoretic solution concept, the Maximum Entropy Blackwell Winner\textit{Maximum Entropy Blackwell Winner} (MaxEntBW\textit{MaxEntBW}), that is well-defined under multi-objective intransitive preferences. To enable computing MaxEntBWs at scale, we derive PROSPER\texttt{PROSPER}: a provably efficient PFT algorithm. Unlike prior self-play techniques, PROSPER\texttt{PROSPER} directly handles multiple objectives without requiring scalarization. We then apply PROSPER\texttt{PROSPER} to the problem of fine-tuning large language models (LLMs) from multi-objective LLM-as-a-Judge feedback (e.g., rubric-based judges), a setting where both sources of intransitivity arise. We find that PROSPER\texttt{PROSPER} outperforms all baselines considered across both instruction following and general chat benchmarks, releasing trained model checkpoints at the 7B and 3B parameter scales.

Keywords

Cite

@article{arxiv.2602.19041,
  title  = {Back to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning},
  author = {Jiahao Zhang and Lujing Zhang and Keltin Grimes and Zhuohao Yu and Gokul Swamy and Zhiwei Steven Wu},
  journal= {arXiv preprint arXiv:2602.19041},
  year   = {2026}
}

Comments

24 pages, 5 figures

R2 v1 2026-07-01T10:46:02.908Z