Preference Models assume Proportional Hazards of Utilities
Machine Learning
2025-08-20 v1 Artificial Intelligence
Machine Learning
Abstract
Approaches for estimating preferences from human annotated data typically involves inducing a distribution over a ranked list of choices such as the Plackett-Luce model. Indeed, modern AI alignment tools such as Reward Modelling and Direct Preference Optimization are based on the statistical assumptions posed by the Plackett-Luce model. In this paper, I will connect the Plackett-Luce model to another classical and well known statistical model, the Cox Proportional Hazards model and attempt to shed some light on the implications of the connection therein.
Keywords
Cite
@article{arxiv.2508.13189,
title = {Preference Models assume Proportional Hazards of Utilities},
author = {Chirag Nagpal},
journal= {arXiv preprint arXiv:2508.13189},
year = {2025}
}