English
Related papers

Related papers: Equivalence Test in Multi-dimensional Space with A…

200 papers

The Binary Space Partitioning-Tree~(BSP-Tree) process was recently proposed as an efficient strategy for space partitioning tasks. Because it uses more than one dimension to partition the space, the BSP-Tree Process is more efficient and…

Machine Learning · Statistics 2020-03-03 Xuhui Fan , Bin Li , Scott A. Sisson

Empirical likelihood enables a nonparametric, likelihood-driven style of inference without restrictive assumptions routinely made in parametric models. We develop a framework for applying empirical likelihood to the analysis of experimental…

Methodology · Statistics 2023-11-08 Eunseop Kim , Steven N. MacEachern , Mario Peruggia

Two semimetrics on probability distributions are proposed, given as the sum of differences of expectations of analytic functions evaluated at spatial or frequency locations (i.e, features). The features are chosen so as to maximize the…

Machine Learning · Statistics 2016-10-31 Wittawat Jitkrittum , Zoltan Szabo , Kacper Chwialkowski , Arthur Gretton

We investigate the statistical task of closeness (or equivalence) testing for multidimensional distributions. Specifically, given sample access to two unknown distributions $\mathbf p, \mathbf q$ on $\mathbb R^d$, we want to distinguish…

Data Structures and Algorithms · Computer Science 2023-11-23 Ilias Diakonikolas , Daniel M. Kane , Sihan Liu

We study how to verify specific frequency distributions when we observe a stream of $N$ data items taken from a universe of $n$ distinct items. We introduce the \emph{relative Fr\'echet distance} to compare two frequency functions in a…

Data Structures and Algorithms · Computer Science 2025-08-26 Claire Mathieu , Michel de Rougemont

We present an efficient method to estimate cross-validation bandwidth parameters for kernel density estimation in very large datasets where ordinary cross-validation is rendered highly inefficient, both statistically and computationally.…

Methodology · Statistics 2016-09-02 Anirban Bhattacharya , Jeffrey D. Hart

This article concerns tests for the two-sample location problem when the dimension is larger than the sample size. The traditional multivariate-rank-based procedures cannot be used in high dimensional settings because the sample scatter…

Methodology · Statistics 2015-06-30 Long Feng

Design of experiments and estimation of treatment effects in large-scale networks, in the presence of strong interference, is a challenging and important problem. Most existing methods' performance deteriorates as the density of the network…

Methodology · Statistics 2020-12-15 Preetam Nandy , Kinjal Basu , Shaunak Chatterjee , Ye Tu

Even though a train/test split of the dataset randomly performed is a common practice, could not always be the best approach for estimating performance generalization under some scenarios. The fact is that the usual machine learning…

Machine Learning · Computer Science 2022-09-09 Carlos Catania , Jorge Guerra , Juan Manuel Romero , Gabriel Caffaratti , Martin Marchetta

This paper presents a procedure for testing the hypothesis that the underlying distribution of the data is elliptical when using robust location and scatter estimators instead of the sample mean and covariance matrix. Under mild assumptions…

Methodology · Statistics 2015-02-20 Ana M. Bianco , Graciela Boente , Isabel M. Rodrigues

The energy test is a powerful binning-free, multi-dimensional and distribution-free tool that can be applied to compare a measurement to a given prediction (goodness-of-fit) or to check whether two data samples originate from the same…

Data Analysis, Statistics and Probability · Physics 2018-04-30 G. Zech

In data centers, tasks are dispatched to various servers to evenly distribute the workload. When a data center considers implementing a new scheduling algorithm, it typically conducts an A/B test prior to deployment to assess the real-world…

Methodology · Statistics 2026-05-29 Nanshan Jia , Ramesh Johari , Nian Si , Zeyu Zheng

The log-normal distribution is one of the most common distributions used for modeling skewed and positive data. It frequently arises in many disciplines of science, specially in the biological and medical sciences. The statistical analysis…

Methodology · Statistics 2020-01-01 Ayanendranath Basu , Abhijit Mandal , Nirian Martin , Leandro Pardo

Based on the test for equality of quantiles originally introduced by Kosorok (1999), we propose new power formulas for the comparison of one quantile between two treatment groups, as well as for the comparison of a collection of quantiles.…

Methodology · Statistics 2026-03-10 Beatriz Farah , Olivier Bouaziz , Aurélien Latouche

Sampling from multivariate normal distributions, subjected to a variety of restrictions, is a problem that is recurrent in statistics and computing. In the present work, we demonstrate a general framework to efficiently sample a…

A relevant question when analyzing spatial point patterns is that of spatial randomness. More specifically, before any model can be fit to a point pattern a first step is to test the data for departures from complete spatial randomness…

Other Statistics · Statistics 2025-04-07 Vaidehi Dixit , Christopher K. Wikle , Scott H. Holan

Consider a big data multiple testing task, where, due to storage and computational bottlenecks, one is given a very large collection of p-values by splitting into manageable chunks and distributing over thousands of computer nodes. This…

Methodology · Statistics 2018-05-08 Subhadeep Mukhopadhyay

Online controlled experiments, now commonly known as A/B testing, are crucial to causal inference and data driven decision making in many internet based businesses. While a simple comparison between a treatment (the feature under test) and…

Applications · Statistics 2015-01-05 Yu Guo , Alex Deng

In modern data analysis, statistical efficiency improvement is expected via effective collaboration among multiple data holders with non-shared data. In this article, we propose a collaborative score-type test (CST) for testing linear…

Methodology · Statistics 2025-04-30 Yifan Gu , Hanfang Yang , Songshan Yang , Hui Zou

In this article, an adaption of an algorithm for the creation of experimental designs by Lekivetz and Jones (2015) is suggested, dealing with constraints around randomization. Split-plot design of experiments is used, when the levels of…

Methodology · Statistics 2020-03-24 Thomas Muehlenstaedt , Maria Lanzerath