English
Related papers

Related papers: Ranking Abuse via Strategic Pairwise Data Perturba…

200 papers

Ranking problems based on pairwise comparisons, such as those arising in online gaming, often involve a large pool of items to order. In these situations, the gap in performance between any two items can be significant, and the smallest and…

Statistics Theory · Mathematics 2022-06-16 Heejong Bong , Alessandro Rinaldo

This technical report studies the problem of ranking from pairwise comparisons in the classical Bradley-Terry-Luce (BTL) model, with a focus on score estimation. For general graphs, we show that, with sufficiently many samples, maximum…

Machine Learning · Statistics 2023-04-17 Yanxi Chen

For nonbalanced paired comparisons, a wide variety of ranking methods have been proposed. One of the best popular methods is the Bradley-Terry model in which the ranking of a set of objects is decided by the maximum likelihood estimates…

Methodology · Statistics 2016-11-07 Ting Yan

Rank aggregation with pairwise comparisons is widely encountered in sociology, politics, economics, psychology, sports, etc . Given the enormous social impact and the consequent incentives, the potential adversary has a strong motivation to…

Artificial Intelligence · Computer Science 2024-07-03 Ke Ma , Qianqian Xu , Jinshan Zeng , Wei Liu , Xiaochun Cao , Yingfei Sun , Qingming Huang

A preference order or ranking aggregated from pairwise comparison data is commonly understood as a strict total order. However, in real-world scenarios, some items are intrinsically ambiguous in comparisons, which may very well be an…

Machine Learning · Computer Science 2018-07-31 Qianqian Xu , Jiechao Xiong , Xinwei Sun , Zhiyong Yang , Xiaochun Cao , Qingming Huang , Yuan Yao

A number of applications (e.g., AI bot tournaments, sports, peer grading, crowdsourcing) use pairwise comparison data and the Bradley-Terry-Luce (BTL) model to evaluate a given collection of items (e.g., bots, teams, students, search…

Machine Learning · Computer Science 2019-06-12 Jingyan Wang , Nihar B. Shah , R. Ravi

The increasing integration of Large Language Model (LLM) based search engines has transformed the landscape of information retrieval. However, these systems are vulnerable to adversarial attacks, especially ranking manipulation attacks,…

Computation and Language · Computer Science 2025-05-19 Xiyang Hu

Benchmarking outcomes increasingly govern trust, selection, and deployment of LLMs, yet these evaluations remain vulnerable to semantically equivalent adversarial perturbations. Prior work on adversarial robustness in NLP has emphasized…

Machine Learning · Computer Science 2025-10-16 Ivan Dubrovsky , Anastasia Orlova , Illarion Iov , Nina Gubina , Irena Gureeva , Alexey Zaytsev

Evaluation leaderboards such as LMArena play a central role in benchmarking large language models by aggregating pairwise human preferences into model rankings, yet the robustness of these rankings remains poorly understood. We present a…

Machine Learning · Computer Science 2026-05-18 Hosna Oyarhoseini , Jimmy Lin , Amir-Hossein Karimi

This paper is concerned with the problem of top-$K$ ranking from pairwise comparisons. Given a collection of $n$ items and a few pairwise comparisons across them, one wishes to identify the set of $K$ items that receive the highest ranks.…

Machine Learning · Statistics 2019-06-13 Yuxin Chen , Jianqing Fan , Cong Ma , Kaizheng Wang

The standard way to evaluate language models on subjective tasks is through pairwise comparisons: an annotator chooses the "better" of two responses to a prompt. Leaderboards aggregate these comparisons into a single Bradley-Terry (BT)…

Machine Learning · Computer Science 2026-02-26 Hadi Khalaf , Serena L. Wang , Daniel Halpern , Itai Shapira , Flavio du Pin Calmon , Ariel D. Procaccia

In this paper, we study a popular method for inference of the Bradley-Terry model parameters, namely the MM algorithm, for maximum likelihood estimation and maximum a posteriori probability estimation. This class of models includes the…

Machine Learning · Statistics 2020-12-29 Milan Vojnovic , Seyoung Yun , Kaifang Zhou

Preference-based data often appear complex and noisy but may conceal underlying homogeneous structures. This paper introduces a novel framework of ranking structure recognition for preference-based data. We first develop an approach to…

Machine Learning · Statistics 2025-11-11 Nan Lu , Jian Shi , Xin-Yu Tian

The task of ranking individuals or teams, based on a set of comparisons between pairs, arises in various contexts, including sporting competitions and the analysis of dominance hierarchies among animals and humans. Given data on which…

Machine Learning · Statistics 2022-10-21 M. E. J. Newman

We propose a method for evaluating the robustness of widely used LLM ranking systems -- variants of a Bradley--Terry model -- to dropping a worst-case very small fraction of preference data. Our approach is computationally fast and easy to…

Machine Learning · Statistics 2026-03-06 Jenny Y. Huang , Yunyi Shen , Dennis Wei , Tamara Broderick

We consider the problem of ranking $n$ players from partial pairwise comparison data under the Bradley-Terry-Luce model. For the first time in the literature, the minimax rate of this ranking problem is derived with respect to the Kendall's…

Statistics Theory · Mathematics 2021-01-22 Pinhan Chen , Chao Gao , Anderson Y. Zhang

Deciding which large language model (LLM) to use is a complex challenge. Pairwise ranking has emerged as a new method for evaluating human preferences for LLMs. This approach entails humans evaluating pairs of model outputs based on a…

Computation and Language · Computer Science 2025-02-18 Roland Daynauth , Christopher Clarke , Krisztian Flautner , Lingjia Tang , Jason Mars

This paper explores the preference-based top-$K$ rank aggregation problem. Suppose that a collection of items is repeatedly compared in pairs, and one wishes to recover a consistent ordering that emphasizes the top-$K$ ranked items, based…

Machine Learning · Computer Science 2015-05-29 Yuxin Chen , Changho Suh

Rankings derived from pairwise comparisons are central to many economic and computational systems. In the context of large language models (LLMs), rankings are typically constructed from human preference data and presented as leaderboards…

Computation and Language · Computer Science 2026-03-05 Angel Rodrigo Avelar Menendez , Yufeng Liu , Xiaowu Dai

Traditional AI safety evaluations on isolated LLMs are insufficient as multi-agent AI ensembles become prevalent, introducing novel emergent risks. This paper introduces the Multi-Agent Emergent Behavior Evaluation (MAEBE) framework to…

Multiagent Systems · Computer Science 2025-07-11 Sinem Erisken , Timothy Gothard , Martin Leitgab , Ram Potham
‹ Prev 1 2 3 10 Next ›