English
Related papers

Related papers: Ranking Metrics: Extending Acceptability and Perfo…

200 papers

Benchmarking has long served as a foundational practice in machine learning and, increasingly, in modern AI systems such as large language models, where shared tasks, metrics, and leaderboards offer a common basis for measuring progress and…

Artificial Intelligence · Computer Science 2026-02-16 Philip Waggoner

We provide a new characterization of second-order stochastic dominance, also known as increasing concave order. The result has an intuitive interpretation that adding a risk with negative expected value in adverse scenarios makes the…

Risk Management · Quantitative Finance 2024-09-30 Yuanying Guan , Muqiao Huang , Ruodu Wang

Ranking is used for a wide array of problems, most notably information retrieval (search). There are a number of popular approaches to the evaluation of ranking such as Kendall's $\tau$, Average Precision, and nDCG. When dealing with…

Information Retrieval · Computer Science 2026-05-08 Denys Katerenchuk , Andrew Rosenberg

We study the problem of aggregating individual preferences over alternatives into a collective ranking. A distinctive feature of our setting is that agents are matched to alternatives. Applications include rankings of colleges or academic…

Theoretical Economics · Economics 2026-02-11 Gaurab Aryal , Thayer Morrill , Peter Troyan

Benchmarks shape scientific conclusions about model capabilities and steer model development. This creates a feedback loop: stronger benchmarks drive better models, and better models demand more discriminative benchmarks. Ensuring benchmark…

Computation and Language · Computer Science 2025-10-01 Arda Uzunoglu , Tianjian Li , Daniel Khashabi

Algorithmic risk assessments are used to inform decisions in a wide variety of high-stakes settings. Often multiple predictive models deliver similar overall performance but differ markedly in their predictions for individual cases, an…

Machine Learning · Computer Science 2021-05-04 Amanda Coston , Ashesh Rambachan , Alexandra Chouldechova

Recently, $\alpha$-Rank, a graph-based algorithm, has been proposed as a solution to ranking joint policy profiles in large scale multi-agent systems. $\alpha$-Rank claimed tractability through a polynomial time implementation with respect…

Multiagent Systems · Computer Science 2020-03-04 Yaodong Yang , Rasul Tutunov , Phu Sakulwongtana , Haitham Bou Ammar

In many applications such as rationing medical care and supplies, university admissions, and the assignment of public housing, the decision of who receives an allocation can be justified by various normative criteria. Such settings have…

Computer Science and Game Theory · Computer Science 2023-05-30 Siddhartha Banerjee , Matthew Eichhorn , David Kempe

This document is an evaluation of the original "Rank-N-Contrast" (arXiv:2210.01189v2) paper published in 2023. This evaluation is done for academic purposes. Deep regression models often fail to capture the continuous nature of sample…

Machine Learning · Computer Science 2025-06-24 Valentin Six , Alexandre Chidiac , Arkin Worlikar

The field of portfolio selection is an active research topic, which combines elements and methodologies from various fields, such as optimization, decision analysis, risk management, data science, forecasting, etc. The modeling and…

Portfolio Management · Quantitative Finance 2020-10-28 A. Georgantas

Embedding data into vector spaces is a very popular strategy of pattern recognition methods. When distances between embeddings are quantized, performance metrics become ambiguous. In this paper, we present an analysis of the ambiguity…

Computer Vision and Pattern Recognition · Computer Science 2019-02-21 Anguelos Nicolaou , Sounak Dey , Vincent Christlein , Andreas Maier , Dimosthenis Karatzas

Reward modeling is central to alignment pipelines such as RLHF, RLAIF, and PPO-based policy optimization, yet its reliability is constrained by limited and heterogeneous human preference data that are expensive to collect at scale. While…

Machine Learning · Computer Science 2026-05-26 Payel Bhattacharjee , Osvaldo Simeone , Ravi Tandon

Nowadays, several crowdsourcing projects exploit social choice methods for computing an aggregate ranking of alternatives given individual rankings provided by workers. Motivated by such systems, we consider a setting where each worker is…

Computer Science and Game Theory · Computer Science 2018-11-27 Ioannis Caragiannis , Xenophon Chatzigeorgiou , George A. Krimpas , Alexandros A. Voudouris

We introduce the concept of \emph{expected exposure} as the average attention ranked items receive from users over repeated samples of the same query. Furthermore, we advocate for the adoption of the principle of equal expected exposure:…

Information Retrieval · Computer Science 2020-10-22 Fernando Diaz , Bhaskar Mitra , Michael D. Ekstrand , Asia J. Biega , Ben Carterette

Benchmarking is a common method for evaluating trajectory prediction models for autonomous driving. Existing benchmarks rely on datasets, which are biased towards more common scenarios, such as cruising, and distance-based metrics that are…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Changhe Chen , Mozhgan Pourkeshavarz , Amir Rasouli

Sparsity or complexity? In modern high-dimensional asset pricing, these are often viewed as competing principles: richer feature spaces appear to favor complexity, while economic intuition has long favored parsimony. We show that this…

General Finance · Quantitative Finance 2026-04-21 Nima Afsharhajari , Jonathan Yu-Meng Li

Metric elicitation is a recent framework for eliciting classification performance metrics that best reflect implicit user preferences based on the task and context. However, available elicitation strategies have been limited to linear (or…

Machine Learning · Statistics 2022-08-23 Gaurush Hiranandani , Jatin Mathur , Harikrishna Narasimhan , Oluwasanmi Koyejo

Rankings are the primary interface through which many online platforms match users to items (e.g. news, products, music, video). In these two-sided markets, not only the users draw utility from the rankings, but the rankings also determine…

Information Retrieval · Computer Science 2020-06-01 Marco Morik , Ashudeep Singh , Jessica Hong , Thorsten Joachims

We provide a constructive way of defining new elicitable risk measures that are characterised by a multiplicative scoring function. We show that depending on the choice of the scoring function's components, the resulting risk measure…

Mathematical Finance · Quantitative Finance 2025-03-06 Akif Ince , Marlon Moresco , Ilaria Peri , Silvana M. Pesenti

Although originally developed to evaluate sets of items, recall is often used to evaluate rankings of items, including those produced by recommender, retrieval, and other machine learning systems. The application of recall without a formal…

Information Retrieval · Computer Science 2024-12-03 Fernando Diaz , Michael D. Ekstrand , Bhaskar Mitra