English
Related papers

Related papers: Accuracy, Repeatability, and Reproducibility of Fi…

200 papers

This study investigates the reliability and validity of five advanced Large Language Models (LLMs), Claude 3.5, DeepSeek v2, Gemini 2.5, GPT-4, and Mistral 24B, for automated essay scoring in a real world higher education context. A total…

Computers and Society · Computer Science 2025-08-05 Andrea Gaggioli , Giuseppe Casaburi , Leonardo Ercolani , Francesco Collova' , Pietro Torre , Fabrizio Davide

There are over 55 different ways to construct a confidence respectively credible interval (CI) for the binomial proportion. Methods to compare them are necessary to decide which should be used in practice. The interval score has been…

Methodology · Statistics 2022-07-08 Lisa J. Hofer , Leonhard Held

Research artifacts are widely shared to support reproducibility, and artifact evaluation (AE) has become common at many leading conferences. However, AE mainly checks whether artifacts work as claimed and can be reproduced. It largely…

Cryptography and Security · Computer Science 2026-05-08 Nanda Rani , Christian Rossow

Probabilistic classifiers output confidence scores along with their predictions, and these confidence scores should be calibrated, i.e., they should reflect the reliability of the prediction. Confidence scores that minimize standard metrics…

Adversarial attack perturbs an image with an imperceptible noise, leading to incorrect model prediction. Recently, a few works showed inherent bias associated with such attack (robustness bias), where certain subgroups in a dataset (e.g.…

Computer Vision and Pattern Recognition · Computer Science 2022-05-06 Gaurav Kumar Nayak , Ruchit Rawal , Rohit Lal , Himanshu Patil , Anirban Chakraborty

Genome-wide association studies (GWAS) are widely used to discover genetic variants associated with diseases. To control false positives, all findings from GWAS need to be verified with additional evidences, even for associations discovered…

Genomics · Quantitative Biology 2026-03-12 Wei Jiang , Jing-Hao Xue , Weichuan Yu

Large-scale simultaneous hypothesis testing appears in many areas such as microarray studies, genome-wide association studies, brain imaging, disease mapping and astronomical surveys. A well-known inference method is to control the false…

Methodology · Statistics 2025-07-22 Xiaoqing Niu , Pengfei Li , Yuejiao Fu

The stochastic multi-armed bandit model is a simple abstraction that has proven useful in many different contexts in statistics and machine learning. Whereas the achievable limit in terms of regret minimization is now well known, our aim is…

Machine Learning · Statistics 2016-11-15 Emilie Kaufmann , Olivier Cappé , Aurélien Garivier

Large Language Models (LLMs) show promise for automated grading, but their outputs can be unreliable. Rather than improving grading accuracy directly, we address a complementary problem: \textit{predicting when an LLM grader is likely to be…

Computation and Language · Computer Science 2026-04-01 Robinson Ferrer , Damla Turgut , Zhongzhou Chen , Shashank Sonkar

Bayes' Theorem confers inherent limitations on the accuracy of screening tests as a function of disease prevalence. We have shown in previous work that a testing system can tolerate significant drops in prevalence, up until a certain…

Methodology · Statistics 2020-09-01 Jacques Balayla

Objectives: Estimation of areas under receiver operating characteristic curves (AUCs) and their differences is a key task in diagnostic studies. We aimed to derive, evaluate, and implement simple sample size formulas for such studies with a…

Methodology · Statistics 2022-08-03 Di Shu , Guangyong Zou

Multiple-choice exams are frequently used as an efficient and objective method to assess learning but they are more vulnerable to answer-copying than tests based on open questions. Several statistical tests (known as indices in the…

Statistics Theory · Mathematics 2014-09-29 Mauricio Romero , Alvaro Riascos , Diego Jara

Linear combinations of multinomial probabilities, such as those resulting from contingency tables, are of use when evaluating classification system performance. While large sample inference methods for these combinations exist, small sample…

Methodology · Statistics 2021-04-20 Katherine A. Batterton , Christine M. Schubert , Richard L. Warr

Objectives: Discussions of fairness in criminal justice risk assessments typically lack conceptual precision. Rhetoric too often substitutes for careful analysis. In this paper, we seek to clarify the tradeoffs between different kinds of…

Machine Learning · Statistics 2025-07-25 Richard A. Berk , Hoda Heidari , Shahin Jabbari , Michael Kearns , Aaron Roth

Bayesian bandit algorithms with approximate Bayesian inference have been widely used in real-world applications. However, there is a large discrepancy between the superior practical performance of these approaches and their theoretical…

Machine Learning · Computer Science 2023-11-13 Ziyi Huang , Henry Lam , Amirhossein Meisami , Haofeng Zhang

Current safety evaluations of large language models rely on single-shot testing, implicitly assuming that model responses are deterministic and representative of the model's safety alignment. We challenge this assumption by investigating…

Machine Learning · Computer Science 2025-12-17 Erik Larsen

In an empirical Bayes analysis, we use data from repeated sampling to imitate inferences made by an oracle Bayesian with extensive knowledge of the data-generating distribution. Existing results provide a comprehensive characterization of…

Methodology · Statistics 2021-09-09 Nikolaos Ignatiadis , Stefan Wager

Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation that higher accuracy implies safer behavior. This assumption is incomplete in medicine,…

Experimentation platforms in industry must often deal with customer trust issues. Platforms must prove the validity of their claims as well as catch issues that arise. As a central quantity estimated by experimentation platforms, the…

Methodology · Statistics 2025-11-21 Kedar Karhadkar , Jack Klys , Daniel Ting , Artem Vorozhtsov , Houssam Nassif

Randomized benchmarking (RB) is an efficient and robust method to characterize gate errors in quantum circuits. Averaging over random sequences of gates leads to estimates of gate errors in terms of the average fidelity. These estimates are…

Quantum Physics · Physics 2019-09-11 Jonas Helsen , Joel J. Wallman , Steven T. Flammia , Stephanie Wehner
‹ Prev 1 3 4 5 6 7 10 Next ›