English
Related papers

Related papers: Accuracy, Repeatability, and Reproducibility of Fi…

200 papers

Single-prompt accuracy is the dominant way to benchmark language models, but it can miss reliability failures that matter. We evaluate a 15-model open-weight corpus, with the main reliability analyses focused on 10 instruct models across…

Computation and Language · Computer Science 2026-05-05 Ranit Karmakar , Jayita Chatterjee

Consistently checking the statistical significance of experimental results is the first mandatory step towards reproducible science. This paper presents a hitchhiker's guide to rigorous comparisons of reinforcement learning algorithms.…

Methodology · Statistics 2022-08-30 Cédric Colas , Olivier Sigaud , Pierre-Yves Oudeyer

Measurements are generally collected as unilateral or bilateral data in clinical trials or observational studies. For example, in ophthalmology studies, the primary outcome is often obtained from one eye or both eyes of an individual. In…

Methodology · Statistics 2021-11-01 Kejia Wang , Chang-Xing Ma

We provide new non-asymptotic false discovery proportion (FDP) confidence envelopes in several multiple testing settings relevant for modern high dimensional-data methods. We revisit the multiple testing scenarios considered in the recent…

Statistics Theory · Mathematics 2024-09-18 Iqraa Meah , Gilles Blanchard , Etienne Roquain

I describe a new likelihood technique, based on counts-in-cells statistics, that I use to analyze repeating in the BATSE 1B and 2B catalogues. Using the 1B data, I find that repeating is preferred over non-repeating by 4.3:1 odds, with a…

Astrophysics · Physics 2009-10-28 Jean M. Quashnock

The field of psychological sciences has been grappling with the replicability crisis. Various issues have been identified as potential sources of this problem. We bring to light a potential source that has largely been overlooked and…

Methodology · Statistics 2025-04-28 Yoav Zeevi , Sofi Astashenko , Liad Mudrik , Yoav Benjamini

Diagnostic accuracy studies assess sensitivity and specificity of a new index test in relation to an established comparator or the reference standard. The development and selection of the index test is usually assumed to be conducted prior…

Methodology · Statistics 2022-08-30 Max Westphal , Antonia Zapf

Confidence calibration of classification models is a technique to estimate the true posterior probability of the predicted class, which is critical for ensuring reliable decision-making in practical applications. Existing confidence…

Methodology · Statistics 2025-02-19 Jinzong Dong , Zhaohui Jiang , Dong Pan , Haoyang Yu

Background: Test suites are frequently used to quantify relevant software attributes, such as quality or productivity. Problem: We have detected that the same response variable, measured using different test suites, yields different…

Software Engineering · Computer Science 2022-04-26 Oscar Dieste , Fernando Uyaguari , Sira Vegas , Natalia Juristo

Biometric recognition is used across a variety of applications from cyber security to border security. Recent research has focused on ensuring biometric performance (false negatives and false positives) is fair across demographic groups.…

Methodology · Statistics 2022-08-24 Michael Schuckers , Sandip Purnapatra , Kaniz Fatima , Daqing Hou , Stephanie Schuckers

In research on eye movements in reading, it is common to analyze a number of canonical dependent measures to study how the effects of a manipulation unfold over time. Although this gives rise to the well-known multiple comparisons problem,…

Applications · Statistics 2016-10-11 Titus von der Malsburg , Bernhard Angele

Complex scientific models where the likelihood cannot be evaluated present a challenge for statistical inference. Over the past two decades, a wide range of algorithms have been proposed for learning parameters in computationally feasible…

Computation · Statistics 2021-12-16 Aden Forrow , Ruth E. Baker

We use the exact finite sample likelihood and statistical decision theory to answer questions of ``why?'' and ``what should you have done?'' using data from randomized experiments and a utility function that prioritizes safety over…

Econometrics · Economics 2024-07-26 Neil Christy , A. E. Kowalski

We study the problem of testing the goodness of fit of categorical count data to a Poisson distribution uniform over the categories, against a class of alternatives defined by excluding an $\ell_p$ ball, $p \leq 2$, of radius $\epsilon$…

Statistics Theory · Mathematics 2025-12-16 Alon Kipnis

In a fixed-confidence pure exploration problem in stochastic multi-armed bandits, an algorithm iteratively samples arms and should stop as early as possible and return the correct answer to a query about the arms distributions. We are…

Machine Learning · Computer Science 2025-02-04 Adrienne Tuynman , Rémy Degenne

Criminal recidivism models are tools that have gained widespread adoption by parole boards across the United States to assist with parole decisions. These models take in large amounts of data about an individual and then predict whether an…

Computers and Society · Computer Science 2022-09-29 Eric Ingram , Furkan Gursoy , Ioannis A. Kakadiaris

In recent years we have seen rapid and significant progress in automatic image description but what are the open problems in this area? Most work has been evaluated using text-based similarity metrics, which only indicate that there have…

Computation and Language · Computer Science 2017-04-14 Emiel van Miltenburg , Desmond Elliott

To make informative public policy decisions in battling the ongoing COVID-19 pandemic, it is important to know the disease prevalence in a population. There are two intertwined difficulties in estimating this prevalence based on testing…

Methodology · Statistics 2020-12-01 Bryan Cai , John P. A. Ioannidis , Eran Bendavid , Lu Tian

The sample mean is among the most well studied estimators in statistics, having many desirable properties such as unbiasedness and consistency. However, when analyzing data collected using a multi-armed bandit (MAB) experiment, the sample…

Statistics Theory · Mathematics 2021-05-03 Jaehyeok Shin , Aaditya Ramdas , Alessandro Rinaldo

Recent work on chain-of-thought (CoT) faithfulness reports single aggregate numbers (e.g., DeepSeek-R1 acknowledges hints 39% of the time), implying that faithfulness is an objective, measurable property of a model. This paper provides…

Computation and Language · Computer Science 2026-03-25 Richard J. Young