Related papers: The false positive risk: a proposal concerning wha…
Recent research has generated hope that inference scaling, such as resampling solutions until they pass verifiers like unit tests, could allow weaker models to match stronger ones. Beyond inference, this approach also enables training…
The widely claimed replicability crisis in science may lead to revised standards of significance. The customary frequentist confidence intervals, calibrated through hypothetical repetitions of the experiment that is supposed to have…
We seek to understand the probability an individual benefits from treatment (PIBT), an inestimable quantity that must be bounded in practice. Given the innate uncertainty in the population-level bounds on PIBT, we seek to better understand…
\citet{Rosenbaum83ps} introduced the notion of the propensity score and discussed its central role in causal inference with observational studies. Their paper, however, caused a fundamental incoherence with an early paper by…
When conducting large scale inference, such as genome-wide association studies or image analysis, nominal $p$-values are often adjusted to improve control over the family-wise error rate (FWER). When the majority of tests are null,…
Many genomic experiments, notably microarray experiments seeking to detect differential gene expression, involve calculating a large number of p-values. This leads to the multiple testing problem: when the number of null hypotheses is…
Hypothesis testing is an essential statistical method in psychology and the cognitive sciences. The problems of traditional null hypothesis significance testing (NHST) have been discussed widely, and among the proposed solutions to the…
Reproducibility, the ability to recompute results, and replicability, the chances other experimenters will achieve a consistent result, are two foundational characteristics of successful scientific research. Consistent findings from…
Evaluation of counterfactual queries (e.g., "If A were true, would C have been true?") is important to fault diagnosis, planning, and determination of liability. In this paper we present methods for computing the probabilities of such…
Statistical significance of both the original and the replication study is a commonly used criterion to assess replication attempts, also known as the two-trials rule in drug development. However, replication studies are sometimes conducted…
We shall show in this paper that there are experiments which are Bernoulli trials with success probability p > 0.5, and which have the curious feature that it is possible to correctly predict the outcome with probability > p.
In multiple hypothesis testing, the volume of data, defined as the number of replications per null times the total number of nulls, usually defines the amount of resource required. On the other hand, power is an important measure of…
Classical probability theory supports probability measures, assigning a fixed positive real value to each event, these measures are far from satisfactory in formulating real-life occurrences. The main innovation of this paper is the…
In a recent simulation study, Goodman et al. (2019) compare several methods with regard to their type I and type II error rates in case of a thick null hypothesis that includes all values that are practically equivalent to the point null…
It is demonstrated that the statistical method of the famous Aspect - Bell experiment requires negative probability densities and negative probabilities from "the thing" researched, else that thing doesn't exist. The thing refers here to…
A standard practice in statistical hypothesis testing is to mention the p-value alongside the accept/reject decision. We show the advantages of mentioning an e-value instead. With p-values, it is not clear how to use an extreme observation…
There are two distinct definitions of 'P-value' for evaluating a proposed hypothesis or model for the process generating an observed dataset. The original definition starts with a measure of the divergence of the dataset from what was…
Many practical studies rely on hypothesis testing procedures applied to data sets with missing information. An important part of the analysis is to determine the impact of the missing data on the performance of the test, and this can be…
In this paper, we draw attention to a problem that is often overlooked or ignored by companies practicing hypothesis testing (A/B testing) in online environments. We show that conducting experiments on limited inventory that is shared…
This contribution to the debate on confidence limits focuses mostly on the case of measurements with `open likelihood', in the sense that it is defined in the text. I will show that, though a prior-free assessment of {\it confidence} is, in…