English

Preference Models assume Proportional Hazards of Utilities

Machine Learning 2025-08-20 v1 Artificial Intelligence Machine Learning

Abstract

Approaches for estimating preferences from human annotated data typically involves inducing a distribution over a ranked list of choices such as the Plackett-Luce model. Indeed, modern AI alignment tools such as Reward Modelling and Direct Preference Optimization are based on the statistical assumptions posed by the Plackett-Luce model. In this paper, I will connect the Plackett-Luce model to another classical and well known statistical model, the Cox Proportional Hazards model and attempt to shed some light on the implications of the connection therein.

Keywords

Cite

@article{arxiv.2508.13189,
  title  = {Preference Models assume Proportional Hazards of Utilities},
  author = {Chirag Nagpal},
  journal= {arXiv preprint arXiv:2508.13189},
  year   = {2025}
}