English
Related papers

Related papers: Data-Driven Nonparametric Existence and Associatio…

200 papers

We introduce a broadly applicable statistical procedure for testing which parametric distribution family generated a random sample of data. The method, termed the Difference in Differential Entropy (DDE) test, provides a unified framework…

Econometrics · Economics 2025-12-15 Ron Mittelhammer , George Judge , Miguel Henry

It is crucial to detect when an instance lies downright too far from the training samples for the machine learning model to be trusted, a challenge known as out-of-distribution (OOD) detection. For neural networks, one approach to this task…

We consider the problem of constructing sequential power-one tests where the null and alternative classes are specified indirectly through historical or offline data. More specifically, given an offline dataset consisting of observations…

Statistics Theory · Mathematics 2026-03-23 Chia-Yu Hsu , Shubhanshu Shekhar

A two-terminal distributed binary hypothesis testing problem over a noisy channel is studied. The two terminals, called the observer and the decision maker, each has access to $n$ independent and identically distributed samples, denoted by…

Other Statistics · Statistics 2023-02-07 Sreejith Sreekumar , Deniz Gündüz

Global data association is an essential prerequisite for robot operation in environments seen at different times or by different robots. Repetitive or symmetric data creates significant challenges for existing methods, which typically rely…

Robotics · Computer Science 2025-09-22 Yixuan Jia , Mason B. Peterson , Qingyuan Li , Yulun Tian , Jonathan P. How

Active learning can reduce the number of samples needed to perform a hypothesis test and to estimate the parameters of a model. In this paper, we revisit the work of Chernoff that described an asymptotically optimal algorithm for performing…

Machine Learning · Statistics 2022-03-14 Subhojyoti Mukherjee , Ardhendu Tripathy , Robert Nowak

We study distribution testing in the standard access model and the conditional access model when the memory available to the testing algorithm is bounded. In both scenarios, the samples appear in an online fashion and the goal is to test…

Data Structures and Algorithms · Computer Science 2023-09-08 Sampriti Roy , Yadu Vasudev

Two-sample feature selection is the problem of finding features that describe a difference between two probability distributions, which is a ubiquitous problem in both scientific and engineering studies. However, existing methods have…

We investigate distribution testing with access to non-adaptive conditional samples. In the conditional sampling model, the algorithm is given the following access to a distribution: it submits a query set $S$ to an oracle, which returns a…

Data Structures and Algorithms · Computer Science 2018-11-06 Gautam Kamath , Christos Tzamos

This paper addresses the problem of fitting a known distribution to the innovation distribution in a class of stationary and ergodic time series models. The asymptotic null distribution of the usual Kolmogorov--Smirnov test based on the…

Statistics Theory · Mathematics 2007-06-13 Hira L. Koul , Shiqing Ling

We characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test…

Information Theory · Computer Science 2021-02-08 Shengyu Zhu , Biao Chen , Zhitang Chen , Pengfei Yang

In multiple change-point problems, different data segments often follow different distributions, for which the changes may occur in the mean, scale or the entire distribution from one segment to another. Without the need to know the number…

Statistics Theory · Mathematics 2014-05-29 Changliang Zou , Guosheng Yin , Long Feng , Zhaojun Wang

Given samples from two non-negative random variables, we propose a family of tests for the null hypothesis that one random variable stochastically dominates the other at the second order. Test statistics are obtained as functionals of the…

Statistics Theory · Mathematics 2023-10-16 Tommaso Lando , Sirio Legramanti

We consider the problem of sequentially testing a simple null hypothesis versus a composite alternative hypothesis that consists of a finite set of densities. We study sequential tests that are based on thresholding of mixture-based…

Statistics Theory · Mathematics 2013-01-23 Georgios Fellouris , Alexander G. Tartakovsky

We point out necessary and sufficient conditions of uniform consistency of nonparametric sets of alternatives for widespread nonparametric tests. Nonparametric sets of alternatives can be defined both in terms of distribution function and…

Statistics Theory · Mathematics 2020-09-01 Mikhail Ermakov

We study a hypothesis testing problem in which data is compressed distributively and sent to a detector that seeks to decide between two possible distributions for the data. The aim is to characterize all achievable encoding rates and…

Information Theory · Computer Science 2011-02-01 Md. Saifur Rahman , Aaron B. Wagner

Given n observations, we study the consistency of a batch of k new observations, in terms of their distribution function. We propose a non-parametric, non-likelihood test based on Edgeworth expansion of the distribution function. The…

Statistics Theory · Mathematics 2009-06-08 Mahendra Mariadassou , Avner Bar-Hen

Nonparametric two-sample testing is a classical problem in inferential statistics. While modern two-sample tests, such as the edge count test and its variants, can handle multivariate and non-Euclidean data, contemporary gargantuan datasets…

Methodology · Statistics 2023-04-28 Trambak Banerjee , Bhaswar B. Bhattacharya , Gourab Mukherjee

We study the problems of sequential nonparametric two-sample and independence testing. Sequential tests process data online and allow using observed data to decide whether to stop and reject the null hypothesis or to collect more data,…

Machine Learning · Statistics 2023-07-21 Aleksandr Podkopaev , Aaditya Ramdas

A massive dataset often consists of a growing number of (potentially) heterogeneous sub-populations. This paper is concerned about testing various forms of heterogeneity arising from massive data. In a general nonparametric framework, a set…

Statistics Theory · Mathematics 2016-01-26 Junwei Lu , Guang Cheng , Han Liu