English
Related papers

Related papers: Generic E-Variables for Exact Sequential k-Sample …

200 papers

Two-sample tests evaluate whether two samples are realizations of the same distribution (the null hypothesis) or two different distributions (the alternative hypothesis). We consider a new setting for this problem where sample features are…

Machine Learning · Computer Science 2022-07-20 Weizhi Li , Gautam Dasarathy , Karthikeyan Natesan Ramamurthy , Visar Berisha

In most prediction and estimation situations, scientists consider various statistical models for the same problem, and naturally want to select amongst the best. Hansen et al. (2011) provide a powerful solution to this problem by the…

Methodology · Statistics 2026-01-23 Sebastian Arnold , Georgios Gavrilopoulos , Benedikt Schulz , Johanna Ziegel

Many real-world applications of flow-based generative models desire a diverse set of samples that cover multiple modes of the target distribution. However, the predominant approach for obtaining diverse sets is not sample-efficient, as it…

Machine Learning · Computer Science 2025-04-11 Mashrur M. Morshed , Vishnu Boddeti

The entropy of an ergodic finite-alphabet process can be computed from a single typical sample path x_1^n using the entropy of the k-block empirical probability and letting k grow with $n$ roughly like log n. We further assume that the…

Probability · Mathematics 2009-11-10 J. -R. Chazottes , D. Gabrielli

Random effects are a flexible addition to statistical models to capture structural heterogeneity in the data, such as spatial dependencies, individual differences, temporal dependencies, or non-linear effects. Testing for the presence (or…

Methodology · Statistics 2024-10-21 Fabio Vieira , Hongwei Zhao , Joris Mulder

We consider sequential hypothesis testing between two quantum states using adaptive and non-adaptive strategies. In this setting, samples of an unknown state are requested sequentially and a decision to either continue or to accept one of…

Quantum Physics · Physics 2023-03-07 Yonglong Li , Vincent Y. F. Tan , Marco Tomamichel

The growing availability of network data and of scientific interest in distributed systems has led to the rapid development of statistical models of network structure. Typically, however, these are models for the entire network, while the…

Statistics Theory · Mathematics 2022-03-18 Cosma Rohilla Shalizi , Alessandro Rinaldo

Persistent entropy (PE) is an information-theoretic summary statistic of persistence barcodes that has been widely used to detect regime changes in complex systems. Despite its empirical success, a general theoretical understanding of when…

Machine Learning · Statistics 2026-02-11 Matteo Rucco

We investigate the optimization of two probabilistic generative models with binary latent variables using a novel variational EM approach. The approach distinguishes itself from previous variational approaches by using latent states as…

Machine Learning · Statistics 2018-02-26 Jörg Lücke , Zhenwen Dai , Georgios Exarchakis

Adjusting for covariates is a well established method to estimate the total causal effect of an exposure variable on an outcome of interest. Depending on the causal structure of the mechanism under study there may be different adjustment…

Statistics Theory · Mathematics 2021-04-27 Jack Kuipers , Giusi Moffa

For a random variable $N = 0, 1, 2, \ldots$ we study the following question: When does the sum of $N$ many independent and identically distributed copies of a random variable $X$ have the same law a a nontrivial rescaling of $X$? We show…

Probability · Mathematics 2026-04-03 Andrey Sarantsev

A word-valued source $\mathbf{Y} = Y_1,Y_2,...$ is discrete random process that is formed by sequentially encoding the symbols of a random process $\mathbf{X} = X_1,X_2,...$ with codewords from a codebook $\mathscr{C}$. These processes…

Information Theory · Computer Science 2009-04-27 Roy Timo , Kim Blackmore , Leif Hanlen

There are two distinct definitions of 'P-value' for evaluating a proposed hypothesis or model for the process generating an observed dataset. The original definition starts with a measure of the divergence of the dataset from what was…

Other Statistics · Statistics 2023-09-25 Sander Greenland

In many environments, only a relatively small subset of the complete state space is necessary in order to accomplish a given task. We develop a simple technique using emergency stops (e-stops) to exploit this phenomenon. Using e-stops…

Machine Learning · Computer Science 2019-12-05 Samuel Ainsworth , Matt Barnes , Siddhartha Srinivasa

We analyze the extreme value dependence of independent, not necessarily identically distributed multivariate regularly varying random vectors. More specifically, we propose estimators of the spectral measure locally at some time point and…

Statistics Theory · Mathematics 2023-06-05 Holger Drees

We propose a novel family of test statistics to detect the presence of changepoints in a sequence of dependent, possibly multivariate, functional-valued observations. Our approach allows to test for a very general class of changepoints,…

Methodology · Statistics 2023-10-10 B. Cooper Boniece , Lajos Horváth , Lorenzo Trapani

Variable selection plays a fundamental role in high-dimensional data analysis. Various methods have been developed for variable selection in recent years. Well-known examples are forward stepwise regression (FSR) and least angle regression…

Methodology · Statistics 2018-02-01 Siliang Gong , Kai Zhang , Yufeng Liu

Two-sample tests for multivariate data and non-Euclidean data are widely used in many fields. Parametric tests are mostly restrained to certain types of data that meets the assumptions of the parametric models. In this paper, we study a…

Methodology · Statistics 2018-05-01 Hao Chen , Xu Chen , Yi Su

Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…

Methodology · Statistics 2024-09-05 F. Richard Guo , Rajen D. Shah

We study the problems of sequential nonparametric two-sample and independence testing. Sequential tests process data online and allow using observed data to decide whether to stop and reject the null hypothesis or to collect more data,…

Machine Learning · Statistics 2023-07-21 Aleksandr Podkopaev , Aaditya Ramdas