English
Related papers

Related papers: Improving Pearson's chi-squared test: hypothesis t…

200 papers

We propose a new approach to sequential testing which is an adaptive (on-line) extension of the (off-line) framework developed in [10]. It relies upon testing of pairs of hypotheses in the case where each hypothesis states that the vector…

Statistics Theory · Mathematics 2017-02-27 Anatoli Juditsky , Arkadi Nemirovski

In the "correlated sampling" problem, two players are given probability distributions $P$ and $Q$, respectively, over the same finite set, with access to shared randomness. Without any communication, the two players are each required to…

Computational Complexity · Computer Science 2020-11-24 Mohammad Bavarian , Badih Ghazi , Elad Haramaty , Pritish Kamath , Ronald L. Rivest , Madhu Sudan

We consider goodness-of-fit tests with i.i.d. samples generated from a categorical distribution $(p_1,...,p_k)$. For a given $(q_1,...,q_k)$, we test the null hypothesis whether $p_j=q_{\pi(j)}$ for some label permutation $\pi$. The…

Statistics Theory · Mathematics 2018-07-30 Chao Gao

Empirical likelihood is a popular nonparametric or semi-parametric statistical method with many nice statistical properties. Yet when the sample size is small, or the dimension of the accompanying estimating function is high, the…

Statistics Theory · Mathematics 2010-10-05 Yukun Liu , Jiahua Chen

In this work, we revisit the one- and two-sample testing problems: binary hypothesis testing in which one or both distributions are unknown. For the one-sample test, we provide a more streamlined proof of the asymptotic optimality of…

Information Theory · Computer Science 2026-04-21 Arick Grootveld , Biao Chen , Venkata Gandikota

Covariate adjustment is an important tool in the analysis of randomized clinical trials and observational studies. It can be used to increase efficiency and thus power, and to reduce possible bias. While most statistical tests in randomized…

Methodology · Statistics 2011-08-03 Xiaoru Wu , Zhiliang Ying

We study the problem of testing the goodness of fit of categorical count data to a Poisson distribution uniform over the categories, against a class of alternatives defined by excluding an $\ell_p$ ball, $p \leq 2$, of radius $\epsilon$…

Statistics Theory · Mathematics 2025-12-16 Alon Kipnis

The statistical analysis of discrete data has been the subject of extensive statistical research dating back to the work of Pearson. In this survey we review some recently developed methods for testing hypotheses about high-dimensional…

Machine Learning · Statistics 2017-12-19 Sivaraman Balakrishnan , Larry Wasserman

The problem of fitting an event distribution when the total expected number of events is not fixed, keeps appearing in experimental studies. In a chi-square fit, if overall normalization is one of the parameters parameters to be fit, the…

Data Analysis, Statistics and Probability · Physics 2015-07-01 Byron Roe

Consider a binary statistical hypothesis testing problem, where $n$ independent and identically distributed random variables $Z^n$ are either distributed according to the null hypothesis $P$ or the alternative hypothesis $Q$, and only $P$…

Information Theory · Computer Science 2024-04-15 K. V. Harsha , Jithin Ravi , Tobias Koch

This paper develops the process of using Richardson Extrapolation to improve the Kernel Density Estimation method, resulting in a more accurate (lower Mean Squared Error) estimate of a probability density function for a distribution of data…

Probability · Mathematics 2018-12-21 Ruben G. Ascoli

Given N data points drawn from a chi-square distribution, we use Bayesian inference to determine most likely values and N-dependent confidence intervals for the width sigma and the number k of degrees of freedom of that distribution. Using…

Nuclear Theory · Physics 2021-06-14 H. -L. Harney , H. A. Weidenmüller

The $k$-of-$n$ testing problem involves performing $n$ independent tests sequentially, in order to determine whether/not at least $k$ tests pass. The objective is to minimize the expected cost of testing. This is a fundamental and…

Data Structures and Algorithms · Computer Science 2026-03-26 Rayen Tan , Viswanath Nagarajan

I present here a generalization of the maximum likelihood method and the $\chi^2$ method to the cases in which the data are {\it not} assumed to be Gaussian distributed. The method, based on the multivariate Edgeworth expansion, can find…

Astrophysics · Physics 2007-05-23 Luca Amendola

We study confidence regions and approximate chi-squared tests for variable groups in high-dimensional linear regression. When the size of the group is small, low-dimensional projection estimators for individual coefficients can be directly…

Statistics Theory · Mathematics 2016-02-23 Ritwik Mitra , Cun-Hui Zhang

Two-sample hypothesis testing-determining whether two sets of data are drawn from the same distribution-is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing,…

Machine Learning · Statistics 2026-04-21 Antoine Chatalic , Marco Letizia , Nicolas Schreuder , Lorenzo Rosasco

Symmetry plays a central role in the sciences, machine learning, and statistics. For situations in which data are known to obey a symmetry, a multitude of methods that exploit symmetry have been developed. Statistical tests for the presence…

Methodology · Statistics 2024-12-24 Kenny Chiu , Benjamin Bloem-Reddy

We propose a new test to address the nonparametric Behrens-Fisher problem involving different distribution functions in the two samples. Our procedure tests the null hypothesis $\mathcal{H}_0: \theta = \frac{1}{2}$, where $\theta = P(X<Y) +…

Methodology · Statistics 2025-04-09 Stephen Schüürhuis , Frank Konietschke , Edgar Brunner

Many statistical hypotheses can be formulated in terms of polynomial equalities and inequalities in the unknown parameters and thus correspond to semi-algebraic subsets of the parameter space. We consider large sample asymptotics for the…

Statistics Theory · Mathematics 2009-04-03 Mathias Drton

Despite many applications, dimensionality reduction in the $\ell_1$-norm is much less understood than in the Euclidean norm. We give two new oblivious dimensionality reduction techniques for the $\ell_1$-norm which improve exponentially…

Data Structures and Algorithms · Computer Science 2021-08-09 Yi Li , David P. Woodruff , Taisuke Yasuda