English
Related papers

Related papers: Confidence and discoveries with e-values

200 papers

While generative models, especially large language models (LLMs), are ubiquitous in today's world, principled mechanisms to assess their (in)correctness are limited. Using the conformal prediction framework, previous works construct sets of…

Machine Learning · Statistics 2026-04-02 Guneet S. Dhillon , Javier González , Teodora Pandeva , Alicia Curth

False discovery rate (FDR) has been widely used as an error measure in large scale multiple testing problems, but most research in the area has been focused on procedures for controlling the FDR based on independent test statistics or the…

Methodology · Statistics 2009-09-29 Weihua Tang , Cun-Hui Zhang

Conformal Prediction (CP) serves as a robust framework that quantifies uncertainty in predictions made by Machine Learning (ML) models. Unlike traditional point predictors, CP generates statistically valid prediction regions, also known as…

Machine Learning · Computer Science 2024-03-29 A. A. Balinsky , A. D. Balinsky

Causal discovery methods based on the PC algorithm are proven to be sound if all structural assumptions are fulfilled and all conditional independence tests are correct. This idealized setting is rarely given in real data. In this work, we…

Machine Learning · Statistics 2026-03-19 Sofia Faltenbacher , Jonas Wahl , Rebecca Herman , Jakob Runge

Hypothesis testing is a central statistical method in psychological research and the cognitive sciences. While the problems of null hypothesis significance testing (NHST) have been debated widely, few attractive alternatives exist. In this…

Methodology · Statistics 2020-06-08 Riko Kelter , Julio Michael Stern

Experimentation platforms in industry must often deal with customer trust issues. Platforms must prove the validity of their claims as well as catch issues that arise. As a central quantity estimated by experimentation platforms, the…

Methodology · Statistics 2025-11-21 Kedar Karhadkar , Jack Klys , Daniel Ting , Artem Vorozhtsov , Houssam Nassif

Adaptive clinical trials rely on interim analyses, flexible stopping, and data-dependent design modifications that complicate statistical guarantees when fixed-horizon test statistics are repeatedly inspected or reused after adaptations.…

Methodology · Statistics 2026-02-09 Alexandra Sokolova , Vadim Sokolov

We analyze different types of simulations that applied researchers can use to assess whether their inference methods reliably control false-positive rates. We show that different assessments involve trade-offs, varying in the types of…

Econometrics · Economics 2025-10-03 Bruno Ferman

Reliability is probability of success in a success-failure experiment. Confidence in reliability estimate improves with increasing number of samples. Assurance sets confidence level same as reliability to create one number for easier…

Methodology · Statistics 2023-03-07 Sanjay M. Joshi

Structural equation models and Bayesian networks have been widely used to study causal relationships between continuous variables. Recently, a non-Gaussian method called LiNGAM was proposed to discover such causal models and has been…

Machine Learning · Statistics 2010-06-23 Yusuke Komatsu , Shohei Shimizu , Hidetoshi Shimodaira

A/B testing is ubiquitous within the machine learning and data science operations of internet companies. Generically, the idea is to perform a statistical test of the hypothesis that a new feature is better than the existing platform---for…

Statistics Theory · Mathematics 2017-10-11 David Goldberg , James E. Johndrow

E-values and E-processes (nonnegative supermartingales) provide anytime-valid evidence for sequential testing via Ville's inequality, yet their connection to Bayesian reasoning, representational structure, and computational feasibility are…

Statistics Theory · Mathematics 2026-03-11 Nicholas G. Polson , Vadim Sokolov , Daniel Zantedeschi

In this article, we derive and compare methods to derive \textit{p}-values and sets of confidence intervals with strong control of the family-wise error rates and coverage for estimates of treatment effects in cluster randomised trials with…

Methodology · Statistics 2023-02-08 Samuel I Watson , Joshua Akinyemi , Karla Hemming

When interpreting A/B tests, we typically focus only on the statistically significant results and take them by face value. This practice, termed post-selection inference in the statistical literature, may negatively affect both point…

Applications · Statistics 2021-06-01 Alex Deng , Yicheng Li , Jiannan Lu , Vivek Ramamurthy

The problem of combining p-values is an old and fundamental one, and the classic assumption of independence is often violated or unverifiable in many applications. There are many well-known rules that can combine a set of arbitrarily…

Statistics Theory · Mathematics 2025-03-21 Matteo Gasparin , Ruodu Wang , Aaditya Ramdas

The closure principle is a standard tool for achieving strong family-wise error rate (FWER) control in multiple testing problems. We develop an e-value-based closed testing framework that inherits nice properties of e-values, which are…

Methodology · Statistics 2026-05-19 Will Hartog , Lihua Lei

Two procedures for checking Bayesian models are compared using a simple test problem based on the local Hubble expansion. Over four orders of magnitude, p-values derived from a global goodness-of-fit criterion for posterior probability…

Instrumentation and Methods for Astrophysics · Physics 2018-06-27 Leon B. Lucy

We address the problem of testing conditional mean and conditional variance for non-stationary data. We build e-values and p-values for four types of non-parametric composite hypotheses with specified mean and variance as well as other…

Statistics Theory · Mathematics 2024-09-25 Yixuan Fan , Zhanyi Jiao , Ruodu Wang

Prediction-powered inference is a recent methodology for the safe use of black-box ML models to impute missing data, strengthening inference of statistical parameters. However, many applications require strong properties besides valid…

We present a novel and easy-to-use method for calibrating error-rate based confidence intervals to evidence-based support intervals. Support intervals are obtained from inverting Bayes factors based on a parameter estimate and its standard…

Methodology · Statistics 2023-06-28 Samuel Pawel , Alexander Ly , Eric-Jan Wagenmakers