English
Related papers

Related papers: Sum-Based Scoring for Dichotomous and Likert-scale…

200 papers

As large language models take on growing roles as automated evaluators in practical settings, a critical question arises: Can individuals persuade an LLM judge to assign unfairly high scores? This study is the first to reveal that…

Computation and Language · Computer Science 2025-08-12 Yerin Hwang , Dongryeol Lee , Taegwan Kang , Yongil Kim , Kyomin Jung

This work studies the applicability of expensive external oracles such as large language models in answering top-k queries over predicted scores. Such scores are incurred by user-defined functions to answer personalized queries over…

Databases · Computer Science 2025-02-19 Sohrab Namazi Nia , Subhodeep Ghosh , Senjuti Basu Roy , Sihem Amer-Yahia

Ranking is used in sport leagues to determine a champion and/or to decide on promotion/relegation of teams. Arguably, the best known ranking method relies on scores obtained by cumulating the points associated with the wins and the draws of…

Physics and Society · Physics 2023-03-28 Leszek Szczecinski

Order statistics provide an intuition for combining multiple lists of scores over a common index set. This intuition is particularly valuable when the lists to be combined cannot be directly compared in a sensible way. We describe here the…

Machine Learning · Computer Science 2020-06-19 Arvind Thiagarajan

This paper proposes matrix-scaled consensus algorithm, which generalizes the scaled consensus algorithm in \cite{Roy2015scaled}. In (scalar) scaled consensus algorithms, the agents' states do not converge to a common value, but to different…

Optimization and Control · Mathematics 2022-08-16 Minh Hoang Trinh , Dung Van Vu , Quoc Van Tran , Hyo-Sung Ahn

Ranking methods or models based on their performance is of prime importance but is tricky because performance is fundamentally multidimensional. In the case of classification, precision and recall are scores with probabilistic…

Performance · Computer Science 2026-03-31 Sébastien Piérard , Adrien Deliège , Marc Van Droogenbroeck

Recently there has been a significant interest in learning disentangled representations, as they promise increased interpretability, generalization to unseen scenarios and faster learning on downstream tasks. In this paper, we investigate…

Machine Learning · Computer Science 2019-10-30 Francesco Locatello , Gabriele Abbati , Tom Rainforth , Stefan Bauer , Bernhard Schölkopf , Olivier Bachem

Question answering (QA) tasks have been extensively studied in the field of natural language processing (NLP). Answers to open-ended questions are highly diverse and difficult to quantify, and cannot be simply evaluated as correct or…

Computation and Language · Computer Science 2024-10-03 Xiaotian Lu , Jiyi Li , Koh Takeuchi , Hisashi Kashima

When comparing performance (of products, services, entities, etc.), multiple attributes are involved. This paper deals with a way of weighting these attributes when one is seeking an overall score. It presents an objective approach to…

Methodology · Statistics 2024-09-04 Chris Tofallis

We provide methods to validate and compare sensor outputs, or inference algorithms applied to sensor data, by adapting statistical scoring rules. The reported output should either be in the form of a prediction interval or of a parameter…

Data Analysis, Statistics and Probability · Physics 2015-07-07 A. D. Martin , T. C. A. Molteno , M. Parry

To overcome the limitations of automated metrics (e.g. BLEU, METEOR) for evaluating dialogue systems, researchers typically use human judgments to provide convergent evidence. While it has been demonstrated that human judgments can suffer…

Computation and Language · Computer Science 2019-09-24 Sashank Santhanam , Samira Shaikh

A system of nested dichotomies is a method of decomposing a multi-class problem into a collection of binary problems. Such a system recursively splits the set of classes into two subsets, and trains a binary classifier to distinguish…

Machine Learning · Statistics 2016-07-06 Tim Leathart , Bernhard Pfahringer , Eibe Frank

We study the supremum of some random Dirichlet polynomials with independent coefficients and obtain sharp upper and lower bounds for supremum expectation thus extending the results from our previous work (see…

Probability · Mathematics 2009-04-23 Mikhail Lifshits , Michel Weber

In contemporary applied and computational mathematics, a frequent challenge is to bound the expectation of the spectral norm of a sum of independent random matrices. This quantity is controlled by the norm of the expected square of the…

Probability · Mathematics 2015-10-19 Joel A. Tropp

One of the most widely used tasks for evaluating Large Language Models (LLMs) is Multiple-Choice Question Answering (MCQA). While open-ended question answering tasks are more challenging to evaluate, MCQA tasks are, in principle, easier to…

Computation and Language · Computer Science 2025-06-10 Francesco Maria Molfese , Luca Moroni , Luca Gioffré , Alessandro Scirè , Simone Conia , Roberto Navigli

Advances in automated scoring are closely aligned with advances in machine-learning and natural-language-processing techniques. With recent progress in large language models (LLMs), the use of ChatGPT, Gemini, Claude, and other…

Computation and Language · Computer Science 2025-09-30 Haowei Hua , Hong Jiao , Dan Song

"LLM-as-a-judge," which utilizes large language models (LLMs) as evaluators, has proven effective in many evaluation tasks. However, evaluator LLMs exhibit numerical bias, a phenomenon where certain evaluation scores are generated…

Computation and Language · Computer Science 2026-01-27 Ayako Sato , Hwichan Kim , Zhousi Chen , Masato Mita , Mamoru Komachi

Question answering-based summarization evaluation metrics must automatically determine whether the QA model's prediction is correct or not, a task known as answer verification. In this work, we benchmark the lexical answer verification…

Computation and Language · Computer Science 2022-04-22 Daniel Deutsch , Dan Roth

We investigate the order of the variance of the optimal alignments score of two independent iid binary random words having the same length. The letters are equiprobable, but the scoring function is such that one letter has a larger score…

Probability · Mathematics 2016-06-17 Christian Houdré , Heinrich Matzinger

In this paper, we study the weighted sums of multiple t-values and of multiple t-star values at even arguments. Some general weighted sum formulas are given, where the weight coefficients are given by (symmetric) polynomials of the…

Number Theory · Mathematics 2019-08-09 Zhonghua Li , Ce Xu
‹ Prev 1 4 5 6 7 8 10 Next ›