English
Related papers

Related papers: Ranking Abuse via Strategic Pairwise Data Perturba…

200 papers

Ranking LLMs via pairwise human feedback underpins current leaderboards for open-ended tasks, such as creative writing and problem-solving. We analyze ~89K comparisons in 116 languages from 52 LLMs from Arena, and show that the best-fit…

Machine Learning · Computer Science 2026-05-08 Jai Moondra , Ayela Chughtai , Bhargavi Lanka , Swati Gupta

The Bradley-Terry-Luce (BTL) model is a classic and very popular statistical approach for eliciting a global ranking among a collection of items using pairwise comparison data. In applications in which the comparison outcomes are observed…

Methodology · Statistics 2022-11-30 Wanshan Li , Daren Wang , Alessandro Rinaldo

Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are taken at face value - we show this is a (potentially costly)…

Most decision-making models, including the pairwise comparison method, assume the decision-makers honesty. However, it is easy to imagine a situation where a decision-maker tries to manipulate the ranking results. This paper presents three…

Artificial Intelligence · Computer Science 2024-10-11 Michał Strada , Sebastian Ernst , Jacek Szybowski , Konrad Kułakowski

Rank aggregation with pairwise comparisons has shown promising results in elections, sports competitions, recommendations, and information retrieval. However, little attention has been paid to the security issue of such algorithms, in…

Machine Learning · Computer Science 2022-09-14 Ke Ma , Qianqian Xu , Jinshan Zeng , Guorong Li , Xiaochun Cao , Qingming Huang

Maximum likelihood estimation furnishes powerful insights into voting theory, and the design of voting rules. However the MLE can usually be badly corrupted by a single outlying sample. This means that a single voter or a group of colluding…

Data Structures and Algorithms · Computer Science 2022-07-19 Allen Liu , Ankur Moitra

While machine learning (ML) methods have received a lot of attention in recent years, these methods are primarily for prediction. Empirical researchers conducting policy evaluations are, on the other hand, pre-occupied with causal problems,…

Machine Learning · Statistics 2019-03-04 Noemi Kreif , Karla DiazOrdaz

We explore the top-$K$ rank aggregation problem. Suppose a collection of items is compared in pairs repeatedly, and we aim to recover a consistent ordering that focuses on the top-$K$ ranked items based on partially revealed preference…

Machine Learning · Computer Science 2016-03-15 Minje Jang , Sunghyun Kim , Changho Suh , Sewoong Oh

As pairwise ranking becomes broadly employed for elections, sports competitions, recommendations, and so on, attackers have strong motivation and incentives to manipulate the ranking list. They could inject malicious comparisons into the…

Machine Learning · Computer Science 2021-07-06 Ke Ma , Qianqian Xu , Jinshan Zeng , Xiaochun Cao , Qingming Huang

The Bradley-Terry-Luce (BTL) model is a popular statistical approach for estimating the global ranking of a collection of items using pairwise comparisons. To ensure accurate ranking, it is essential to obtain precise estimates of the model…

Statistics Theory · Mathematics 2022-06-24 Wanshan Li , Shamindra Shrotriya , Alessandro Rinaldo

\textit{Mallows model} is a widely-used probabilistic framework for learning from ranking data, with applications ranging from recommendation systems and voting to aligning language models with human preferences~\cite{chen2024mallows,…

Machine Learning · Statistics 2025-07-14 Yeganeh Alimohammadi , Kiana Asgari

Abstract Like electoral systems, decision-making methods are also vulnerable to manipulation by decision-makers. The ability to effectively defend against such threats can only come from thoroughly understanding the manipulation mechanisms.…

Artificial Intelligence · Computer Science 2024-03-25 Jacek Szybowski , Konrad Kułakowski , Jiri Mazurek , Sebastian Ernst

This paper introduces a simple efficient learning algorithms for general sequential decision making. The algorithm combines Optimism for exploration with Maximum Likelihood Estimation for model estimation, which is thus named OMLE. We prove…

Machine Learning · Computer Science 2022-11-24 Qinghua Liu , Praneeth Netrapalli , Csaba Szepesvári , Chi Jin

The Bradley-Terry model is widely used for the analysis of pairwise comparison data and, in essence, produces a ranking of the items under comparison. We embed the Bradley-Terry model within a stochastic block model, allowing items to…

Methodology · Statistics 2025-11-06 Lapo Santi , Nial Friel

Estimating the mean counterfactual outcome under a treatment rule is a central problem in causal inference and policy evaluation. Standard estimators, including inverse probability weighting (IPW), augmented IPW (AIPW), and targeted maximum…

Methodology · Statistics 2026-05-06 Yichen Xu , Mark J. van der Laan

Principal stratification is a widely used framework for addressing post-randomization complications. After using principal stratification to define causal effects of interest, researchers are increasingly turning to finite mixture models to…

Methodology · Statistics 2019-08-20 Avi Feller , Evan Greif , Nhat Ho , Luke Miratrix , Natesh Pillai

Value alignment, which aims to ensure that large language models (LLMs) and other AI agents behave in accordance with human values, is critical for ensuring safety and trustworthiness of these systems. A key component of value alignment is…

Artificial Intelligence · Computer Science 2025-03-11 Ziwei Xu , Mohan Kankanhalli

This paper considers the problem of ranking objects based on their latent merits using data from pairwise interactions. We allow for incomplete observation of these interactions and study what can be inferred about rankings in such…

Econometrics · Economics 2025-09-23 Federico Crippa , Danil Fedchenko

Large language models (LLMs) are increasingly being deployed in high-stakes applications like hiring, yet their potential for unfair decision-making remains understudied in generative and retrieval settings. In this work, we examine the…

Computation and Language · Computer Science 2025-09-05 Preethi Seshadri , Hongyu Chen , Sameer Singh , Seraphina Goldfarb-Tarrant

We study the ranking of individuals, teams, or objects, based on pairwise comparisons between them, using the Bradley-Terry model. Estimates of rankings within this model are commonly made using a simple iterative algorithm first introduced…

Machine Learning · Statistics 2023-08-16 M. E. J. Newman