English
Related papers

Related papers: The Least Difference in Means: A Statistic for Eff…

200 papers

Despite its importance to experimental design, statistical power (the probability that, given a real effect, an experiment will reject the null hypothesis) has largely been ignored by the NLP community. Underpowered experiments make it more…

Computation and Language · Computer Science 2020-10-15 Dallas Card , Peter Henderson , Urvashi Khandelwal , Robin Jia , Kyle Mahowald , Dan Jurafsky

Empirical science needs to be based on facts and claims that can be reproduced. This calls for replicating the studies that proclaim the claims, but practice in most fields still fails to implement this idea. When such studies emerged in…

Other Statistics · Statistics 2025-08-27 Werner A. Stahel

Real-world data typically contain a large number of features that are often heterogeneous in nature, relevance, and also units of measure. When assessing the similarity between data points, one can build various distance measures using…

Machine Learning · Statistics 2022-05-27 Aldo Glielmo , Claudio Zeni , Bingqing Cheng , Gabor Csanyi , Alessandro Laio

Large language models (LLMs) have enabled a range of applications in zero-shot and few-shot learning settings, including the generation of synthetic datasets for training and testing. However, to reliably use these synthetic datasets, it is…

Computation and Language · Computer Science 2024-09-19 Gaurav Maheshwari , Dmitry Ivanov , Kevin El Haddad

To measure how well pretrained representations encode some linguistic property, it is common to use accuracy of a probe, i.e. a classifier trained to predict the property from the representations. Despite widespread adoption of probes,…

Computation and Language · Computer Science 2020-03-30 Elena Voita , Ivan Titov

Accurate power and sample size (PSS) calculations are essential for designing studies that use quasi-likelihood (QL) models, which extend generalized linear models (GLMs) to settings where the full distribution of the outcome is not…

Methodology · Statistics 2026-03-03 Shijie Yuan , Amy Cochran , Paul Rathouz

For testing the statistical significance of a treatment effect, we usually compare between two parts of a population, one is exposed to the treatment, and the other is not exposed to it. Standard parametric and nonparametric two-sample…

Computation · Statistics 2012-11-02 Bikram Karmakar , Kumaresh Dhara , Kushal Kumar Dey , Analabha Basu , Anil Ghosh

Stratifying factors, like age and gender, can modify the effect of treatments and exposures on risk of a studied outcome. Several effect measures, including the relative risk, hazard ratio, odds ratio, and risk difference, can be used to…

Methodology · Statistics 2021-11-05 Jake Shannin , Babette A. Brumback

We show that publishing results using the statistical significance filter---publishing only when the p-value is less than 0.05---leads to a vicious cycle of overoptimistic expectation of the replicability of results. First, we show…

Methodology · Statistics 2017-05-16 Shravan Vasishth , Andrew Gelman

Governments sometimes need to analyse sets of research papers within a field in order to monitor progress, assess the effect of recent policy changes, or identify areas of excellence. They may compare the average citation impacts of the…

Digital Libraries · Computer Science 2015-09-22 Ruth Fairclough , Mike Thelwall

Two semimetrics on probability distributions are proposed, given as the sum of differences of expectations of analytic functions evaluated at spatial or frequency locations (i.e, features). The features are chosen so as to maximize the…

Machine Learning · Statistics 2016-10-31 Wittawat Jitkrittum , Zoltan Szabo , Kacper Chwialkowski , Arthur Gretton

When designing and evaluating an experiment or observational study, it is useful to have a realistic hypothesis regarding the average treatment effect. We present an approach to conceptualizing this average by first considering a…

Methodology · Statistics 2026-04-10 Andrew Gelman , Amy Krefman , Lauren Kennedy , Jessica Hullman

Meta-analysis employs statistical techniques to synthesize the results of individual studies, providing an estimate of the overall effect size for a specific outcome of interest. The direction and magnitude of this estimate, along with its…

Econometrics · Economics 2025-03-06 Amin Haghnejad , Mahboobeh Farahati

When the individual studies assembled for a meta-analysis report means ($\mu_C$, $\mu_T$) for their treatment (T) and control (C) arms, but those data are on different scales or come from different instruments, the customary measure of…

Methodology · Statistics 2023-04-18 Elena Kulinskaya , David C. Hoaglin

In multigroup data settings with small within-group sample sizes, standard $F$-tests of group-specific linear hypotheses can have low power, particularly if the within-group sample sizes are not large relative to the number of explanatory…

Methodology · Statistics 2022-03-25 Andrew McCormack , Peter Hoff

Significant treatment effects are often emphasized when interpreting and summarizing empirical findings in studies that estimate multiple, possibly many, treatment effects. Under this kind of selective reporting, conventional treatment…

Econometrics · Economics 2025-12-11 Andreas Dzemski , Ryo Okui , Wenjie Wang

Consistently checking the statistical significance of experimental results is the first mandatory step towards reproducible science. This paper presents a hitchhiker's guide to rigorous comparisons of reinforcement learning algorithms.…

Methodology · Statistics 2022-08-30 Cédric Colas , Olivier Sigaud , Pierre-Yves Oudeyer

The last few years have seen the proliferation of measures that quantify the scientific output of researches. Yet, these measures focus on productivity, thus fostering the "publish or perish" paradigm. This article proposes a measure that…

Digital Libraries · Computer Science 2013-04-15 Alejandro M. Aragón

We distinguish two questions (i) how much information does the prior contain? and (ii) what is the effect of the prior? Several measures have been proposed for quantifying effective prior sample size, for example Clarke [1996] and Morita et…

Methodology · Statistics 2020-01-30 David E Jones , Robert N Trangucci , Yang Chen

Hypothesis testing in high dimensional data is a notoriously difficult problem without direct access to competing models' likelihood functions. This paper argues that statistical divergences can be used to quantify the difference between…

Data Analysis, Statistics and Probability · Physics 2024-08-02 Jeremy J. H. Wilkinson , Christopher G. Lester