Related papers: Significance Tests in Climate Science
We briefly review some of the scientific challenges and epistemological issues related to climate science. We discuss the formulation and testing of theories and numerical models, which, given the presence of unavoidable uncertainties in…
We present a general Bayesian method for quantifying the statistical reliability of one-dimensional measures of scientific quality based on citation data. Two quality measures used in practice -- ``papers per year'' and ``Hirsch's $h$'' --…
In response to growing concern about the reliability and reproducibility of published science, researchers have proposed adopting measures of greater statistical stringency, including suggestions to require larger sample sizes and to lower…
Statistical significance testing plays an important role when drawing conclusions from experimental results in NLP papers. Particularly, it is a valuable tool when one would like to establish the superiority of one algorithm over another.…
While running any experiment, we often have to consider the statistical power to ensure an effective study. Statistical power or power ensures that we can observe an effect with high probability if such a true effect exists. However,…
We analyze the notion that physical theories are quantitative and testable by observations in experiments. This leads us to propose a new, Bayesian, interpretation of probabilities in physics that unifies their current use in classical…
A theory of measurement uncertainty is presented, which, since it is based exclusively on the Bayesian approach and on the subjective concept of conditional probability, is applicable in the most general cases. The recent International…
This article, produced as a result of the Symposium on Statistical Inference, is an introduction to the literature on the function of expertise, judgment, and choice in the practice of statistics and scientific research. In particular,…
We point out that the ideas underlying some test procedures recently proposed for testing post-model-selection (and for some other test problems) in the econometrics literature have been around for quite some time in the statistics…
The basic idea of importance sampling is to use independent samples from a proposal measure in order to approximate expectations with respect to a target measure. It is key to understand how many samples are required in order to guarantee…
Measurements play a crucial role in doing physics: Their results provide the basis on which we adopt or reject physical theories. In this note, we examine the effect of subjecting measurements themselves to our experience. We require that…
An exciting recent development is the uptake of deep neural networks in many scientific fields, where the main objective is outcome prediction with the black-box nature. Significance testing is promising to address the black-box issue and…
I compare and discuss critically several measures of statistical significance in common use in astrophysics and in high energy physics. I also exhibit some relationships among them.
Although we accept that Physics is, as a last resort, an experimental science, the relationship between theory and experiment is far away from being trivial. Any experiment is always explained within a determinate theoretical context and,…
Background and objective. Circular statistics and Rayleigh tests are important tools for analyzing the occurrence of cyclic events. However, current methods fail in the presence of measurement bias, such as incomplete or otherwise…
We consider the conditional randomization test as a way to account for covariate imbalance in randomized experiments. The test accounts for covariate imbalance by comparing the observed test statistic to the null distribution of the test…
Empirical science needs to be based on facts and claims that can be reproduced. This calls for replicating the studies that proclaim the claims, but practice in most fields still fails to implement this idea. When such studies emerged in…
Statistical models that include random effects are commonly used to analyze longitudinal and correlated data, often with strong and parametric assumptions about the random effects distribution. There is marked disagreement in the literature…
Measuring performance & quantifying a performance change are core evaluation techniques in programming language and systems research. Of 122 recent scientific papers, as many as 65 included experimental evaluation that quantified a…
With limited resources, scientific inquiries must be prioritized for further study, funding, and translation based on their practical significance: whether the effect size is large enough to be meaningful in the real world. Doing so must…