English
Related papers

Related papers: Accuracy, Repeatability, and Reproducibility of Fi…

200 papers

In statistical inference, confidence set procedures are typically evaluated based on their validity and width properties. Even when procedures achieve rate-optimal widths, confidence sets can still be excessively wide in practice due to…

Statistics Theory · Mathematics 2025-03-20 Kenta Takatsu

Many experiments can be interpreted in terms of random processes operating according to some internal protocols. When experiments are costly or cannot be repeated only one or a few finite samples are available. In this paper we study data…

Data Analysis, Statistics and Probability · Physics 2016-02-02 Marian Kupczynski , Hans De Raedt

In this paper, we study the effects of using an algorithm-based risk assessment instrument to support the prediction of risk of criminalrecidivism. The instrument we use in our experiments is a machine learning version ofRiskEval(name…

Computers and Society · Computer Science 2024-03-18 Manuel Portela , Carlos Castillo , Songül Tolan , Marzieh Karimi-Haghighi , Antonio Andres Pueyo

Analogy-Based Estimation (ABE) is a popular method for non-algorithmic estimation due to its simplicity and effectiveness. The Analogy-Based Estimation (ABE) model was proposed by researchers, however, no optimal approach for reliable…

Software Engineering · Computer Science 2025-12-02 Tarun Chintada , Uday Kiran Cheera

We present a procedure for handling asymmetric errors. Many results in particle physics are presented as values with different positive and negative errors, and there is no consistent procedure for handling them. We consider the difference…

Methodology · Statistics 2025-08-05 Roger Barlow , Alessandra Brazzale , Igor Volobouev

False positive results in screening tests have potentially severe psychological, medical, and financial consequences for the recipient. However, there have been few efforts to quantify how the risk of a false positive accumulates over time.…

Applications · Statistics 2022-06-20 Tim White , Sara Algeri

We propose an empirical likelihood test that is able to test the goodness of fit of a class of parametric and semi-parametric multiresponse regression models. The class includes as special cases fully parametric models; semi-parametric…

Statistics Theory · Mathematics 2010-01-12 Song Xi Chen , Ingrid Van Keilegom

Is there a statistical difference between Naive Bayes and Random Forest in terms of recall, f-measure, and precision for predicting software defects? By utilizing systematic literature review and meta-analysis, we are answering this…

Software Engineering · Computer Science 2025-02-06 Ch Muhammad Awais , Wei Gu , Gcinizwe Dlamini , Zamira Kholmatova , Giancarlo Succi

The Bayes Error Rate (BER) is the fundamental limit on the achievable generalizable classification accuracy of any machine learning model due to inherent uncertainty within the data. BER estimators offer insight into the difficulty of any…

Machine Learning · Computer Science 2025-09-24 Lesley Wheat , Martin v. Mohrenschildt , Saeid Habibi

Pretrial risk assessment tools are used on over one million U.S. defendants each year, yet their use for predicting rare violent re-offense faces a basic statistical barrier. We derive a universal precision bound -- the Likelihood Ratio…

Computers and Society · Computer Science 2026-05-01 Marco Pollanen

The validity screen (Cacioli, 2026d, 2026e) classifies LLM confidence signals as Valid, Indeterminate, or Invalid. We test whether these classifications predict selective prediction performance. Twenty frontier LLMs from seven families were…

Computation and Language · Computer Science 2026-04-21 Jon-Paul Cacioli

Fast Radio Bursts (FRBs), a class of millisecond-scale, highly energetic phenomena with unknown progenitors and radiation mechanisms, require proper statistical analysis as a key method for uncovering their mysteries. In this research, we…

High Energy Astrophysical Phenomena · Physics 2025-02-27 Xianghan Cui , Clancy James , Di Li , Chengmin Zhang

Mixture models are a popular tool in model-based clustering. Such a model is often fitted by a procedure that maximizes the likelihood, such as the EM algorithm. At convergence, the maximum likelihood parameter estimates are typically…

Computation · Statistics 2019-07-23 Adrian O'Hagan , Thomas Brendan Murphy , Luca Scrucca , Isobel Claire Gormley

As large language models (LLMs) are increasingly deployed in critical decision-making systems, the lack of reliable methods to measure their uncertainty presents a fundamental trustworthiness risk. We introduce a normalized confidence score…

Machine Learning · Computer Science 2026-03-10 Xie Xiaohu , Liu Xiaohu , Yao Benjamin

Several application domains require formal but flexible approaches to the comparison problem. Different process models that cannot be related by behavioral equivalences should be compared via a quantitative notion of similarity, which is…

Logic in Computer Science · Computer Science 2010-06-29 Alessandro Aldini

We develop an Empirical Bayes grading scheme that balances the informativeness of the assigned grades against the expected frequency of ranking errors. Applying the method to a massive correspondence experiment, we grade the racial biases…

Econometrics · Economics 2023-06-23 Patrick Kline , Evan K. Rose , Christopher R. Walters

Diagnostic tests are almost never perfect. Studies quantifying their performance use knowledge of the true health status, measured with a reference diagnostic test. Researchers commonly assume that the reference test is perfect, which is…

Applications · Statistics 2024-08-20 Filip Obradović

Reproducibility is essential to reliable scientific discovery in high-throughput experiments. In this work we propose a unified approach to measure the reproducibility of findings identified from replicate experiments and identify putative…

Applications · Statistics 2011-10-24 Qunhua Li , James B. Brown , Haiyan Huang , Peter J. Bickel

Analogues of the frequentist chi-square and F tests are proposed for testing goodness-of-fit and consistency for Bayesian models. Simple examples exhibit these tests' detection of inconsistency between consecutive experiments with identical…

Instrumentation and Methods for Astrophysics · Physics 2016-03-16 L. B. Lucy

We present a violation of the CHSH inequality without the fair sampling assumption with a continuously pumped photon pair source combined with two high efficiency superconducting detectors. Due to the continuous nature of the source, the…

‹ Prev 1 4 5 6 7 8 10 Next ›