English
Related papers

Related papers: When More Is Less: Pitfalls of significance testin…

200 papers

In physics the value of a theory is measured by its agreement with experimental data. But how should the physics community gauge the value of an emerging theory that has not been tested experimentally as of yet? With no reality check, a…

Instrumentation and Methods for Astrophysics · Physics 2011-08-29 Abraham Loeb

Testing hypotheses is an issue of primary importance in the scientific research, as well as in many other human activities. Much clarification about it can be achieved if the process of learning from data is framed in a stochastic model of…

Data Analysis, Statistics and Probability · Physics 2007-05-23 G. D'Agostini

A number of information retrieval studies have been done to assess which statistical techniques are appropriate for comparing systems. However, these studies are focused on TREC-style experiments, which typically have fewer than 100 topics.…

Information Retrieval · Computer Science 2023-05-15 Ngozi Ihemelandu , Michael D. Ekstrand

Hypothesis testing results often rely on simple, yet important assumptions about the behaviour of the distribution of p-values under the null and the alternative. We examine tests for one dimensional parameters of interest that converge to…

Statistics Theory · Mathematics 2021-08-06 Yanbo Tang , Radu Craiu , Lei Sun

Benford's law is often used as a support to critical decisions related to data quality or the presence of data manipulations or even fraud. However, many authors argue that conventional statistical tests will reject the null of data…

Methodology · Statistics 2022-06-16 Roy Cerqueti , Claudio Lupi

The following proposition is justified from several different points of view. If you use P = 0.05 to suggest that you have made a discovery, you will be wrong at least 30 percent of the time. If, as is often the case, experiments are…

Applications · Statistics 2014-11-21 David Colquhoun

When statisticians quarrel about hypothesis testing, the debate usually focus on which method is the correct one. The fundamental question of whether we should test hypothesis at all tends to be forgotten. This lack of debate has its roots…

Other Statistics · Statistics 2016-11-22 André C. R. Martins

Empirical phi-divergence test-statistics have demostrated to be a useful technique for the simple null hypothesis to improve the finite sample behaviour of the classical likelihood ratio test-statistic, as well asfor model misspecification…

Methodology · Statistics 2016-01-15 Narayanaswamy Balakrishnan , Nirian Martin , Leandro Pardo

Score reliability is necessary for establishing a validity argument for an instrument, and is therefore highly important to investigate. Depending on the proposed instrument use and score interpretations, differing degrees of precision in…

Physics Education · Physics 2017-02-23 Robert M. Talbot

Statistical analysis is often used to evaluate the evidence for or against scientific hypotheses, and various statistics (e.g., p-values, likelihood ratios, Bayes factors) are interpreted as measures of evidence strength. Here I consider…

Other Statistics · Statistics 2018-05-30 Veronica J. Vieland

In several large-scale replication projects, statistically non-significant results in both the original and the replication study have been interpreted as a "replication success". Here we discuss the logical problems with this approach:…

Methodology · Statistics 2023-12-19 Samuel Pawel , Rachel Heyard , Charlotte Micheloud , Leonhard Held

There is a general agreement that it is important to consider the practical relevance of an effect in addition to its statistical significance, yet a formal definition of practical relevance is still pending and shall be provided within…

Methodology · Statistics 2021-10-20 Patrick Schwaferts , Thomas Augustin

Verifying that a statistically significant result is scientifically meaningful is not only good scientific practice, it is a natural way to control the Type I error rate. Here we introduce a novel extension of the p-value - a…

Methodology · Statistics 2018-07-04 Jeffrey D. Blume , Lucy DAgostino McGowan , William D. Dupont , Robert A. Greevy

We consider the problem of estimating the number of false null hypotheses among a very large number of independently tested hypotheses, focusing on the situation in which the proportion of false null hypotheses is very small. We propose a…

Statistics Theory · Mathematics 2007-06-13 Nicolai Meinshausen , John Rice

Despite their importance in supporting experimental conclusions, standard statistical tests are often inadequate for research areas, like the life sciences, where the typical sample size is small and the test assumptions difficult to…

Methodology · Statistics 2011-04-15 Pietro Berkes , Jozsef Fiser

Statistical insignificance does not suggest the absence of effect, yet scientists must often use null results as evidence of negligible (near-zero) effect size to falsify scientific hypotheses. Doing so must assess a result's null strength,…

Labelling data is a major practical bottleneck in training and testing classifiers. Given a collection of unlabelled data points, we address how to select which subset to label to best estimate test metrics such as accuracy, $F_1$ score or…

Machine Learning · Computer Science 2021-09-27 Emine Yilmaz , Peter Hayes , Raza Habib , Jordan Burgess , David Barber

The classical theory for the meta-analysis of $p$-values is based on the assumption that if the overall null hypothesis is true, then all $p$-values used in a chosen combined test statistic are genuine, i.e., are observations from…

Computation · Statistics 2024-10-08 Rui Santos , M. Fátima Brilhante , Sandra Mendonça

The score test statistic using the observed information is easy to compute numerically. Its large sample distribution under the null hypothesis is well known and is equivalent to that of the score test based on the expected information, the…

Statistics Theory · Mathematics 2018-08-10 N. Karavarsamis , G. Guillera-Arroita , RM Huggins , B J T Morgan

An association rule is statistically significant, if it has a small probability to occur by chance. It is well-known that the traditional frequency-confidence framework does not produce statistically significant rules. It can both accept…

Databases · Computer Science 2014-05-07 Wilhelmiina Hämäläinen