English
Related papers

Related papers: The false positive risk: a proposal concerning wha…

200 papers

Recent research has generated hope that inference scaling, such as resampling solutions until they pass verifiers like unit tests, could allow weaker models to match stronger ones. Beyond inference, this approach also enables training…

Machine Learning · Computer Science 2026-03-27 Benedikt Stroebl , Sayash Kapoor , Arvind Narayanan

The widely claimed replicability crisis in science may lead to revised standards of significance. The customary frequentist confidence intervals, calibrated through hypothetical repetitions of the experiment that is supposed to have…

Statistics Theory · Mathematics 2020-02-11 Luigi Pace , Alessandra Salvan

We seek to understand the probability an individual benefits from treatment (PIBT), an inestimable quantity that must be bounded in practice. Given the innate uncertainty in the population-level bounds on PIBT, we seek to better understand…

Methodology · Statistics 2024-04-04 Gabriel Ruiz , Oscar Hernan Madrid Padilla

\citet{Rosenbaum83ps} introduced the notion of the propensity score and discussed its central role in causal inference with observational studies. Their paper, however, caused a fundamental incoherence with an early paper by…

Methodology · Statistics 2022-03-29 Peng Ding , Tianyu Guo

When conducting large scale inference, such as genome-wide association studies or image analysis, nominal $p$-values are often adjusted to improve control over the family-wise error rate (FWER). When the majority of tests are null,…

Methodology · Statistics 2017-07-20 Sarah Fletcher Mercaldo , Jeffrey D. Blume

Many genomic experiments, notably microarray experiments seeking to detect differential gene expression, involve calculating a large number of p-values. This leads to the multiple testing problem: when the number of null hypotheses is…

Quantitative Methods · Quantitative Biology 2007-05-23 David R. Bickel

Hypothesis testing is an essential statistical method in psychology and the cognitive sciences. The problems of traditional null hypothesis significance testing (NHST) have been discussed widely, and among the proposed solutions to the…

Methodology · Statistics 2020-05-28 Riko Kelter

Reproducibility, the ability to recompute results, and replicability, the chances other experimenters will achieve a consistent result, are two foundational characteristics of successful scientific research. Consistent findings from…

Applications · Statistics 2015-06-23 Jeffrey T. Leek , Roger D. Peng

Evaluation of counterfactual queries (e.g., "If A were true, would C have been true?") is important to fault diagnosis, planning, and determination of liability. In this paper we present methods for computing the probabilities of such…

Artificial Intelligence · Computer Science 2013-02-28 Alexander Balke , Judea Pearl

Statistical significance of both the original and the replication study is a commonly used criterion to assess replication attempts, also known as the two-trials rule in drug development. However, replication studies are sometimes conducted…

Applications · Statistics 2024-05-31 Leonhard Held , Samuel Pawel , Charlotte Micheloud

We shall show in this paper that there are experiments which are Bernoulli trials with success probability p > 0.5, and which have the curious feature that it is possible to correctly predict the outcome with probability > p.

Other Statistics · Statistics 2018-01-09 James D. Stein

In multiple hypothesis testing, the volume of data, defined as the number of replications per null times the total number of nulls, usually defines the amount of resource required. On the other hand, power is an important measure of…

Statistics Theory · Mathematics 2009-06-05 Zhiyi Chi

Classical probability theory supports probability measures, assigning a fixed positive real value to each event, these measures are far from satisfactory in formulating real-life occurrences. The main innovation of this paper is the…

Probability · Mathematics 2009-02-09 Yehuda Izhakian , Zur Izhakian

In a recent simulation study, Goodman et al. (2019) compare several methods with regard to their type I and type II error rates in case of a thick null hypothesis that includes all values that are practically equivalent to the point null…

Methodology · Statistics 2022-06-07 Robin Tim Dreher , Leona Hoffmann , Arne Kramer-Sunderbrink , Peter Pütz , Robin Werner

It is demonstrated that the statistical method of the famous Aspect - Bell experiment requires negative probability densities and negative probabilities from "the thing" researched, else that thing doesn't exist. The thing refers here to…

General Physics · Physics 2025-05-07 Han Geurdes

A standard practice in statistical hypothesis testing is to mention the p-value alongside the accept/reject decision. We show the advantages of mentioning an e-value instead. With p-values, it is not clear how to use an extreme observation…

Methodology · Statistics 2024-04-04 Peter Grünwald

There are two distinct definitions of 'P-value' for evaluating a proposed hypothesis or model for the process generating an observed dataset. The original definition starts with a measure of the divergence of the dataset from what was…

Other Statistics · Statistics 2023-09-25 Sander Greenland

Many practical studies rely on hypothesis testing procedures applied to data sets with missing information. An important part of the analysis is to determine the impact of the missing data on the performance of the test, and this can be…

Methodology · Statistics 2011-02-15 Dan L. Nicolae , Xiao-Li Meng , Augustine Kong

In this paper, we draw attention to a problem that is often overlooked or ignored by companies practicing hypothesis testing (A/B testing) in online environments. We show that conducting experiments on limited inventory that is shared…

Probability · Mathematics 2020-06-11 Dennis Bohle , Alexander Marynych , Matthias Meiners

This contribution to the debate on confidence limits focuses mostly on the case of measurements with `open likelihood', in the sense that it is defined in the text. I will show that, though a prior-free assessment of {\it confidence} is, in…

High Energy Physics - Experiment · Physics 2007-05-23 G. D'Agostini
‹ Prev 1 4 5 6 7 8 10 Next ›