English
Related papers

Related papers: Improving Pearson's chi-squared test: hypothesis t…

200 papers

We consider linear regression in the high-dimensional regime where the number of observations $n$ is smaller than the number of parameters $p$. A very successful approach in this setting uses $\ell_1$-penalized least squares (a.k.a. the…

Methodology · Statistics 2014-02-05 Adel Javanmard , Andrea Montanari

Parton distributions functions (PDFs), which are essential to the interpretation of data from high energy colliders, are measured by representing them as functional forms containing many parameters. Those parameters are determined by…

High Energy Physics - Phenomenology · Physics 2015-03-13 Jon Pumplin

We propose a goodness-of-fit test for degree-corrected stochastic block models (DCSBM). The test is based on an adjusted chi-square statistic for measuring equality of means among groups of $n$ multinomial distributions with $d_1,\dots,d_n$…

Statistics Theory · Mathematics 2022-09-23 Linfan Zhang , Arash A. Amini

In this paper, we prove a local limit theorem for the chi-square distribution with $r > 0$ degrees of freedom and noncentrality parameter $\lambda \geq 0$. We use it to develop refined normal approximations for the survival function. Our…

Statistics Theory · Mathematics 2022-07-29 Frédéric Ouimet

This paper concerns the development of Stein's method for chi-square approximation and its application to problems in statistics. New bounds for the derivatives of the solution of the gamma Stein equation are obtained. These bounds involve…

Probability · Mathematics 2017-05-30 Robert E. Gaunt , Alastair Pickett , Gesine Reinert

Counting experiments often rely on Monte Carlo simulations for predictions of Poisson expectations. The accompanying uncertainty from the finite Monte Carlo sample size can be incorporated into parameter estimation by modifying the Poisson…

Instrumentation and Methods for Astrophysics · Physics 2020-04-22 Thorsten Glüsenkamp

This paper introduces a new method for testing the statistical significance of estimated parameters in predictive regressions. The approach features a new family of test statistics that are robust to the degree of persistence of the…

Econometrics · Economics 2025-02-04 Jean-Yves Pitarakis

We obtain bounds to quantify the distributional approximation in the delta method for vector statistics (the sample mean of $n$ independent random vectors) for normal and non-normal limits, measured using smooth test functions. For normal…

Statistics Theory · Mathematics 2023-05-11 Robert E. Gaunt , Heather Sutcliffe

The Chernoff information between two probability measures is a statistical divergence measuring their deviation defined as their maximally skewed Bhattacharyya distance. Although the Chernoff information was originally introduced for…

Information Theory · Computer Science 2022-10-04 Frank Nielsen

We consider the problem of closeness testing for two discrete distributions in the practically relevant setting of \emph{unequal} sized samples drawn from each of them. Specifically, given a target error parameter $\varepsilon > 0$, $m_1$…

Machine Learning · Computer Science 2015-04-20 Bhaswar B. Bhattacharya , Gregory Valiant

We consider a problem of simple hypothesis testing using a randomized test via a tunable loss function proposed by Liao \textit{et al}. In this problem, we derive results that correspond to the Neyman--Pearson lemma, the Chernoff--Stein…

Information Theory · Computer Science 2022-08-30 Akira Kamatsuka

We study the problem, introduced by Qiao and Valiant, of learning from untrusted batches. Here, we assume $m$ users, all of whom have samples from some underlying distribution $p$ over $1, \ldots, n$. Each user sends a batch of $k$ i.i.d.…

Data Structures and Algorithms · Computer Science 2019-11-07 Sitan Chen , Jerry Li , Ankur Moitra

We present novel bounds for estimating discrete probability distributions under the $\ell_\infty$ norm. These are nearly optimal in various precise senses, including a kind of instance-optimality. Our data-dependent convergence guarantees…

Statistics Theory · Mathematics 2024-02-14 Aryeh Kontorovich , Amichai Painsky

The likelihood ratio test is widely used in exploratory factor analysis to assess the model fit and determine the number of latent factors. Despite its popularity and clear statistical rationale, researchers have found that when the…

Statistics Theory · Mathematics 2025-01-08 Yinqiu He , Zi Wang , Gongjun Xu

Categorical variables are of uttermost importance in biomedical research. When two of them are considered, it is often the case that one wants to test whether or not they are statistically dependent. We show weaknesses of classical methods…

The C statistics, also known as the Cash statistic, is often used in astronomy for the analysis of low-count Poisson data. One of the challenges of the C statistic is that its probability distribution, under the null hypothesis that the…

High Energy Astrophysical Phenomena · Physics 2019-12-12 M. Bonamente

Wilk's theorem, which offers universal chi-squared approximations for likelihood ratio tests, is widely used in many scientific hypothesis testing problems. For modern datasets with increasing dimension, researchers have found that the…

Statistics Theory · Mathematics 2020-08-14 Yinqiu He , Bo Meng , Zhenghao Zeng , Gongjun Xu

We provide an algorithm for properly learning mixtures of two single-dimensional Gaussians without any separability assumptions. Given $\tilde{O}(1/\varepsilon^2)$ samples from an unknown mixture, our algorithm outputs a mixture that is…

Data Structures and Algorithms · Computer Science 2014-05-20 Constantinos Daskalakis , Gautam Kamath

Evolve and resequence studies provide a popular approach to simulate evolution in the lab and explore its genetic basis. In this context, the chi-square test, Fishers exact test, as well as the Cochran-Mantel-Haenszel test are commonly used…

Applications · Statistics 2019-02-22 Kerstin Spitzer , Marta Pelizzola , Andreas Futschik

A hypothesis testing algorithm is replicable if, when run on two different samples from the same distribution, it produces the same output with high probability. This notion, defined by by Impagliazzo, Lei, Pitassi, and Sorell [STOC'22],…

Data Structures and Algorithms · Computer Science 2025-09-05 Anders Aamand , Maryam Aliakbarpour , Justin Y. Chen , Shyam Narayanan , Sandeep Silwal
‹ Prev 1 3 4 5 6 7 10 Next ›