English
Related papers

Related papers: Sum-Based Scoring for Dichotomous and Likert-scale…

200 papers

Automated answer matching, which leverages LLMs to evaluate free-text responses by comparing them to a reference answer, shows substantial promise as a scalable and aligned alternative to human evaluation. However, its reliability requires…

Computation and Language · Computer Science 2026-01-15 Manas Khatore , Sumana Sridharan , Kevork Sulahian , Benjamin J. Smith , Shi Feng

This paper systematically compares different methods of deriving item-level predictions of language models for multiple-choice tasks. It compares scoring methods for answer options based on free generation of responses, various…

Computation and Language · Computer Science 2024-03-05 Polina Tsvilodub , Hening Wang , Sharon Grosch , Michael Franke

Large language models (LLMs) are increasingly used as automated evaluators, yet prior works demonstrate that these LLM judges often lack consistency in scoring when the prompt is altered. However, the effect of the grading scale itself…

Comparing the top $k$ elements between two or more ranked results is a common task in many contexts and settings. A few measures have been proposed to compare top $k$ lists with attractive mathematical properties, but they face a number of…

Information Theory · Computer Science 2013-10-02 Arun Konagurthu , James Collier

Cross-level interactions among fixed effects in linear mixed models (also known as multilevel models) are often complicated by the variances stemming from random effects and residuals. When these variances change across clusters, tests of…

Methodology · Statistics 2022-03-18 Ting Wang , Edgar C. Merkle , Joaquin A. Anguera , Brandon M. Turner

Content-focused research-based assessment instruments typically use items (i.e., questions) as the unit of assessment for scoring, reporting, and validation. Couplet scoring employs an alternative unit of assessment called a couplet, which…

Physics Education · Physics 2024-12-11 Michael Vignal , Gayle Geschwind , Marcos D. Caballero , H. J. Lewandowski

As large-language models have been increasingly used as automatic raters for evaluating free-form content, including document summarization, dialog, and story generation, work has been dedicated to evaluating such models by measuring their…

Computation and Language · Computer Science 2025-09-09 Logan Lawrence , Ashton Williamson , Alexander Shelton

The statistical leverage scores of a complex matrix $A\in\mathbb{C}^{n\times d}$ record the degree of alignment between col$(A)$ and the coordinate axes in $\mathbb{C}^n$. These score are used in random sampling algorithms for solving…

Machine Learning · Statistics 2016-10-03 James Hook

The score test statistic using the observed information is easy to compute numerically. Its large sample distribution under the null hypothesis is well known and is equivalent to that of the score test based on the expected information, the…

Statistics Theory · Mathematics 2018-08-10 N. Karavarsamis , G. Guillera-Arroita , RM Huggins , B J T Morgan

A neutrosophic set is a more general platform, which can be used to present uncertainty, imprecise, incomplete and inconsistent. In this paper a score function and an accuracy function for single valued neutrosophic sets is firstly proposed…

Artificial Intelligence · Computer Science 2014-12-18 Rıdvan Şahin

When applied to question answering and other text generation tasks, language models (LMs) may be queried generatively (by sampling answers from their output distribution) or discriminatively (by using them to score or rank a set of…

Computer Science and Game Theory · Computer Science 2023-10-16 Athul Paul Jacob , Yikang Shen , Gabriele Farina , Jacob Andreas

We study randomly stopped sums via their asymptotic scales. First, finiteness of moments is considered. To generalise this study, asymptotic scales applicable to the class of all heavy-tailed random variables are used. The stopping is…

Probability · Mathematics 2014-05-12 Jaakko Lehtomaa

This paper proposes a class of origin-smooth approximators of indicators underlying the sum-of-negative-part statistic for testing multiple inequalities. The need for simulation or bootstrap to obtain test critical values is thereby…

Methodology · Statistics 2012-06-27 Le-Yu Chen , Jerzy Szroeter

We introduce a new discrepancy score between two distributions that gives an indication on their similarity. While much research has been done to determine if two samples come from exactly the same distribution, much less research…

Machine Learning · Computer Science 2012-10-16 Maayan Harel , Shie Mannor

As with all measurements, the measurement of examinee ability, in terms of scores that the examinee obtains in a test, is also error-ridden. The quantification of such error or uncertainty in the test score data--or rather the complementary…

Applications · Statistics 2015-03-13 Satyendra Nath Chakrabartty , Kangrui Wang , Dalia Chakrabarty

This paper offers a solution method that allows one to find exact values for a large class of convergent series of rational terms. Sums of this form arise often in problems dealing with Quantum Field Theory.

Mathematical Physics · Physics 2007-05-23 Costas Efthimiou

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which the definition of the system-level correlation is inconsistent…

Computation and Language · Computer Science 2022-04-22 Daniel Deutsch , Rotem Dror , Dan Roth

We present a unified approach which gives completely elementary proofs of three weighted sum formulae for double zeta values. This approach also leads to new evaluations of sums relating to the harmonic numbers, the alternating double zeta…

Number Theory · Mathematics 2012-06-13 James Wan

In order to conduct analyses of networked systems where connections between individuals take on a range of values - counts, continuous strengths or ordinal rankings - a common technique is to dichotomize the data according to their…

Applications · Statistics 2015-03-17 Andrew C. Thomas , Joseph K. Blitzstein

Measurements are generally collected as unilateral or bilateral data in clinical trials or observational studies. For example, in ophthalmologic studies, statistical tests are often based on one or two eyes of an individual. For bilateral…

Methodology · Statistics 2020-10-08 Chang-Xing Ma , Kejia Wang