English
Related papers

Related papers: Accuracy, Repeatability, and Reproducibility of Fi…

200 papers

We consider the evaluation of laboratory practice through the comparison of measurements made by participating metrology laboratories when the measurement procedures are considered to have both fixed effects (the residual error due to…

Applications · Statistics 2012-05-09 John F Clare , Annette Koo , Robert B Davies

Reconstruction of shooting events occasionally requires testing of bullets at velocities significantly below the typical muzzle velocity of cartridge arms. Trajectory, drag, and terminal performance depend strongly on velocity, and…

Medical Physics · Physics 2008-12-31 Michael Courtney , Amy Courtney

Reproducibility is a fundamental requirement for validating scientific claims in computational research. Stochastic computational models are widely used in fields such as systems biology, financial modeling and environmental sciences.…

This paper provides a user's guide to the general theory of approximate randomization tests developed in Canay, Romano, and Shaikh (2017) when specialized to linear regressions with clustered data. An important feature of the methodology is…

Econometrics · Economics 2022-03-16 Yong Cai , Ivan A. Canay , Deborah Kim , Azeem M. Shaikh

The principle of peer review is central to the evaluation of research, by ensuring that only high-quality items are funded or published. But peer review has also received criticism, as the selection of reviewers may introduce biases in the…

Other Statistics · Statistics 2015-07-24 Olivier Francois

Biometric systems are widely used for identity verification and identification, including authentication (i.e., one-to-one matching to verify a claimed identity) and identification (i.e., one-to-many matching to find a subject in a…

Cryptography and Security · Computer Science 2024-12-20 Axel Durbet , Paul-Marie Grollemund , Pascal Lafourcade , Kevin Thiry-Atighehchi

Language models (LMs) should provide reliable confidence estimates to help users detect mistakes in their outputs and defer to human experts when necessary. Asking a language model to assess its confidence ("Score your confidence from…

Computation and Language · Computer Science 2025-02-04 Vaishnavi Shrivastava , Ananya Kumar , Percy Liang

Fast Radio Bursts (FRBs) are highly energetic millisecond-duration astrophysical phenomena typically categorized as repeaters or non-repeaters. However, observational limitations may result in misclassifications, potentially leading to a…

High Energy Astrophysical Phenomena · Physics 2025-04-16 Da-Chun Qiang , Jie Zheng , Zhi-Qiang You , Sheng Yang

Properly benchmarking a system is a difficult and intricate task. Unfortunately, even a seemingly innocuous benchmarking mistake can compromise the guarantees provided by a given systems security defense and also put its reproducibility and…

Cryptography and Security · Computer Science 2018-01-09 Erik van der Kouwe , Dennis Andriesse , Herbert Bos , Cristiano Giuffrida , Gernot Heiser

Background: Estimations of causal effects from observational data are subject to various sources of bias. One method of adjusting for the residual biases in the estimation of a treatment effect is through negative control outcomes, where…

Methodology · Statistics 2022-07-29 Hon Hwang , Juan C Quiroz , Blanca Gallego

While existing benchmarks demonstrate the near-perfect performance of large language models (LLMs) on various tasks, this apparent saturation often obscures the need for rigorous evaluation of their reliability. In real-world deployment,…

Machine Learning · Computer Science 2026-05-13 Eungyeup Kim , Chenchen Gu , Vashisth Tiwari , J. Zico Kolter

A prime goal of quantum tomography is to provide quantitatively rigorous characterisation of quantum systems, be they states, processes or measurements, particularly for the purposes of trouble-shooting and benchmarking experiments in…

Quantum Physics · Physics 2015-06-12 Nathan K. Langford

We consider the problem of finding, through adaptive sampling, which of $n$ options (arms) has the largest mean. Our objective is to determine a rule which identifies the best arm with a fixed minimum confidence using as few observations as…

Machine Learning · Computer Science 2022-03-17 MohammadJavad Azizi , Sheldon M Ross , Zhengyu Zhang

For an AI system to be reliable, the confidence it expresses in its decisions must match its accuracy. To assess the degree of match, examples are typically binned by confidence and the per-bin mean confidence and accuracy are compared.…

Machine Learning · Computer Science 2022-02-14 Rebecca Roelofs , Nicholas Cain , Jonathon Shlens , Michael C. Mozer

Background Study individuals may face repeated events overtime. However, there is no consensus around learning approaches to use in a high-dimensional framework for survival data (when the number of variables exceeds the number of…

Methodology · Statistics 2022-03-30 Juliette Murris , Anais Charles-Nelson , Audrey Lavenu , Sandrine Katsahian

We prove a new version of the quantum threshold theorem that applies to concatenation of a quantum code that corrects only one error, and we use this theorem to derive a rigorous lower bound on the quantum accuracy threshold epsilon_0. Our…

Quantum Physics · Physics 2007-05-23 Panos Aliferis , Daniel Gottesman , John Preskill

In the context of the widely used competing risks set-up we discuss different inference procedures for testing equality of two cumulative incidence functions, where the data may be subject to independent right-censoring or left-truncation.…

Statistics Theory · Mathematics 2015-10-13 Dennis Dobler , Markus Pauly

Many prediction tasks can admit multiple models that can perform almost equally well. This phenomenon can can undermine interpretability and safety when competing models assign conflicting predictions to individuals. In this work, we study…

Machine Learning · Computer Science 2025-08-01 Erin George , Deanna Needell , Berk Ustun

Performance evaluations are critical for quantifying algorithmic advances in reinforcement learning. Recent reproducibility analyses have shown that reported performance results are often inconsistent and difficult to replicate. In this…

Machine Learning · Computer Science 2020-08-14 Scott M. Jordan , Yash Chandak , Daniel Cohen , Mengxue Zhang , Philip S. Thomas

Ensemble models often achieve higher accuracy than single learners, but their ability to maintain small generalization gaps is not always well understood. This study examines how ensembles balance accuracy and overfitting across four…

Machine Learning · Computer Science 2025-12-08 Zubair Ahmed Mohammad