English
Related papers

Related papers: This is not normal! (Re-) Evaluating the lower $n$…

200 papers

The framework of distribution testing is currently ubiquitous in the field of property testing. In this model, the input is a probability distribution accessible via independently drawn samples from an oracle. The testing task is to…

Data Structures and Algorithms · Computer Science 2022-09-22 Sourav Chakraborty , Eldar Fischer , Arijit Ghosh , Gopinath Mishra , Sayantan Sen

Classical two-sample permutation tests for equality of distributions have exact size in finite samples, but they fail to control size for testing equality of parameters that summarize each distribution. This paper proposes permutation tests…

Econometrics · Economics 2022-04-22 Marinho Bertanha , EunYi Chung

In the setting of entangled single-sample distributions, the goal is to estimate some common parameter shared by a family of $n$ distributions, given one single sample from each distribution. This paper studies mean estimation for entangled…

Machine Learning · Computer Science 2020-07-14 Yingyu Liang , Hui Yuan

The sample frequency spectrum (SFS) of DNA sequences from a collection of individuals is a summary statistic which is commonly used for parametric inference in population genetics. Despite the popularity of SFS-based inference methods,…

Populations and Evolution · Quantitative Biology 2015-06-24 Jonathan Terhorst , Yun S. Song

Learning joint probability distributions on n random variables requires exponential sample size in the generic case. Here we consider the case that a temporal (or causal) order of the variables is known and that the (unknown) graph of…

Machine Learning · Computer Science 2007-05-23 Pawel Wocjan , Dominik Janzing , Thomas Beth

Linear regression is a fundamental building block of statistical data analysis. It amounts to estimating the parameters of a linear model that maps input features to corresponding outputs. In the classical setting where the precision of…

Computer Science and Game Theory · Computer Science 2019-12-16 Nicolas Gast , Stratis Ioannidis , Patrick Loiseau , Benjamin Roussillon

Background: Clinical prediction models are increasingly used to inform healthcare decisions, but determining the minimum sample size for their development remains a critical and unresolved challenge. Inadequate sample sizes can lead to…

Machine Learning · Computer Science 2026-03-02 Diana Shamsutdinova , Felix Zimmer , Oyebayo Ridwan Olaniran , Sarah Markham , Daniel Stahl , Gordon Forbes , Ewan Carr

Having a sufficient quantity of quality data is a critical enabler of training effective machine learning models. Being able to effectively determine the adequacy of a dataset prior to training and evaluating a model's performance would be…

Machine Learning · Computer Science 2026-04-28 Arya Hatamian , Lionel Levine , Haniyeh Ehsani Oskouie , Majid Sarrafzadeh

Studies often estimate associations between an outcome and multiple variates. For example, studies of diagnostic test accuracy estimate sensitivity and specificity, and studies of predictive and prognostic factors typically estimate…

This note examines the behavior of generalization capabilities - as defined by out-of-sample mean squared error (MSE) - of Linear Gaussian (with a fixed design matrix) and Linear Least Squares regression. Particularly, we consider a…

Statistics Theory · Mathematics 2021-09-21 Karthik Duraisamy

Imbalanced regression refers to prediction tasks where the target variable is skewed. This skewness hinders machine learning models, especially neural networks, which concentrate on dense regions and therefore perform poorly on…

Machine Learning · Computer Science 2025-08-11 Shayan Alahyari , Mike Domaratzki

Regression adjustment, sometimes known as Controlled-experiment Using Pre-Experiment Data (CUPED), is an important technique in internet experimentation. It decreases the variance of effect size estimates, often cutting confidence interval…

Methodology · Statistics 2023-11-30 Daniel Ting , Kenneth Hung

Time series forecasting is one of the most active research topics. Machine learning methods have been increasingly adopted to solve these predictive tasks. However, in a recent work, these were shown to systematically present a lower…

Machine Learning · Statistics 2019-10-01 Vitor Cerqueira , Luis Torgo , Carlos Soares

Conventionally, regression discontinuity analysis contrasts a univariate regression's limits as its independent variable, $R$, approaches a cut-point, $c$, from either side. Alternative methods target the average treatment effect in a small…

Applications · Statistics 2021-06-21 Adam C. Sales , Ben B. Hansen

Class distribution skews in imbalanced datasets may lead to models with prediction bias towards majority classes, making fair assessment of classifiers a challenging task. Metrics such as Balanced Accuracy are commonly used to evaluate a…

In this paper, we propose a regression model where the response variable is beta prime distributed using a new parameterization of this distribution that is indexed by mean and precision parameters. The proposed regression model is useful…

Methodology · Statistics 2018-04-23 Marcelo Bourguignon , Manoel Santos-Neto , Mário de Castro

We study confidence intervals based on hard-thresholding, soft-thresholding, and adaptive soft-thresholding in a linear regression model where the number of regressors $k$ may depend on and diverge with sample size $n$. In addition to the…

Statistics Theory · Mathematics 2018-10-08 Ulrike Schneider

Let ${X_1,...,X_n}$ be i.i.d. random observations. Let $\mathbb{S}=\mathbb{L}+\mathbb{T}$ be a $U$-statistic of order $k\ge2$ where $\mathbb{L}$ is a linear statistic having asymptotic normal distribution, and $\mathbb{T}$ is a…

Probability · Mathematics 2009-12-14 Vidmantas Bentkus , Bing-Yi Jing , Wang Zhou

Regression is the workhorse of statistics, and is often faced with real data that contain outliers. When these are casewise outliers, that is, cases that are entirely wrong or belong to a different population, the issue can be remedied by…

Methodology · Statistics 2026-03-06 Jakob Raymaekers , Peter J. Rousseeuw

Standard random-effects meta-analysis methods perform poorly when applied to few studies only. Such settings however are commonly encountered in practice. It is unclear, whether or to what extent small-sample-size behaviour can be improved…

Methodology · Statistics 2019-01-15 Svenja E. Seide , Christian Röver , Tim Friede