English
Related papers

Related papers: Efficient and Stable Multi-Dimensional Kolmogorov-…

200 papers

Given an i.i.d. sample drawn from a density $f$, we propose to test that $f$ equals some prescribed density $f_0$ or that $f$ belongs to some translation/scale family. We introduce a multiple testing procedure based on an estimation of the…

Statistics Theory · Mathematics 2016-08-16 Magalie Fromont , Béatrice Laurent

Group-invariant probability distributions appear in many data-generative models in machine learning, such as graphs, point clouds, and images. In practice, one often needs to estimate divergences between such distributions. In this work, we…

Machine Learning · Computer Science 2026-02-05 Behrooz Tahmasebi , Stefanie Jegelka

`Distribution regression' refers to the situation where a response Y depends on a covariate P where P is a probability distribution. The model is Y=f(P) + mu where f is an unknown regression function and mu is a random error. Typically, we…

Machine Learning · Statistics 2013-02-04 Barnabas Poczos , Alessandro Rinaldo , Aarti Singh , Larry Wasserman

We provide improved differentially private algorithms for identity testing of high-dimensional distributions. Specifically, for $d$-dimensional Gaussian distributions with known covariance $\Sigma$, we can test whether the distribution…

Data Structures and Algorithms · Computer Science 2022-07-26 Shyam Narayanan

The traditional lower bound estimation method for powerlaw distributions based on the Kolmogorov-Smirnov distance proved to perform better than other competing methods. However, if applied to very large collections of data, such a method…

Computation · Statistics 2015-03-19 Alessandro Bessi

Consider $n$ iid random variables, where $\xi_1, \ldots, \xi_n$ are $n$ realisations of a random variable $\xi$ and $\zeta_1, \ldots, \zeta_n$ are $n$ realisations of a random variable $\zeta$. The distribution of each realisation of $\xi$,…

Probability · Mathematics 2018-03-06 Tommy Liu

We derive an optimal shrinkage sample covariance matrix (SCM) estimator which is suitable for high dimensional problems and when sampling from an unspecified elliptically symmetric distribution. Specifically, we derive the optimal (oracle)…

Methodology · Statistics 2017-07-03 Esa Ollila

We review a finite-sampling exponential bound due to Serfling and discuss related exponential bounds for the hypergeometric distribution. We then discuss how such bounds motivate some new results for two-sample empirical processes. Our…

Statistics Theory · Mathematics 2017-02-20 Evan Greene , Jon A. Wellner

In this paper, we develop a systematic theory for high dimensional analysis of variance in multivariate linear regression, where the dimension and the number of coefficients can both grow with the sample size. We propose a new \emph{U}~type…

Methodology · Statistics 2023-01-12 Zhipeng Lou , Xianyang Zhang , Wei Biao Wu

The purpose of this paper is twofold. On a technical side, we propose an extension of the Hausdorff distance from metric spaces to spaces equipped with asymmetric distance measures. Specifically, we focus on the family of Bregman…

Machine Learning · Computer Science 2025-04-11 Tuyen Pham , Hana Dal Poz Kouřimská , Hubert Wagner

A limit theorem for the largest interpoint distance of $p$ independent and identically distributed points in $\mathbb{R}^n$ to the Gumbel distribution is proved, where the number of points $p=p_n$ tends to infinity as the dimension of the…

Probability · Mathematics 2024-02-13 Johannes Heiny , Carolin Kleemann

Two semimetrics on probability distributions are proposed, given as the sum of differences of expectations of analytic functions evaluated at spatial or frequency locations (i.e, features). The features are chosen so as to maximize the…

Machine Learning · Statistics 2016-10-31 Wittawat Jitkrittum , Zoltan Szabo , Kacper Chwialkowski , Arthur Gretton

Distances between probability distributions are a key component of many statistical machine learning tasks, from two-sample testing to generative modeling, among others. We introduce a novel distance between measures that compares them…

Machine Learning · Statistics 2025-07-09 Arturo Castellanos , Anna Korba , Pavlo Mozharovskyi , Hicham Janati

Traditional problems in computational geometry involve aspects that are both discrete and continuous. One such example is nearest-neighbor searching, where the input is discrete, but the result depends on distances, which vary continuously.…

Computational Geometry · Computer Science 2023-08-21 Ahmed Abdelkader , David M. Mount

Finite precision approximations of discrete probability distributions are considered, applicable for distribution synthesis, e.g., probabilistic shaping. Two algorithms are presented that find the optimal $M$-type approximation $Q$ of a…

Information Theory · Computer Science 2017-05-08 Georg Böcherer , Bernhard C. Geiger

Optimal Transport (OT) metrics allow for defining discrepancies between two probability measures. Wasserstein distance is for longer the celebrated OT-distance frequently-used in the literature, which seeks probability distributions to be…

Machine Learning · Computer Science 2021-10-14 Mokhtar Z. Alaya , Gilles Gasso , Maxime Berar , Alain Rakotomamonjy

We propose a minimum distance estimation method for robust regression in sparse high-dimensional settings. The traditional likelihood-based estimators lack resilience against outliers, a critical issue when dealing with high-dimensional…

Methodology · Statistics 2013-07-12 Aurélie C. Lozano , Nicolai Meinshausen

Gradients have been exploited in proposal distributions to accelerate the convergence of Markov chain Monte Carlo algorithms on discrete distributions. However, these methods require a natural differentiable extension of the target discrete…

Machine Learning · Computer Science 2023-02-28 Yue Xiang , Dongyao Zhu , Bowen Lei , Dongkuan Xu , Ruqi Zhang

In this paper, we derive new, nearly optimal bounds for the Gaussian approximation to scaled averages of $n$ independent high-dimensional centered random vectors $X_1,\dots,X_n$ over the class of rectangles in the case when the covariance…

Probability · Mathematics 2021-05-13 Victor Chernozhukov , Denis Chetverikov , Yuta Koike

We want to approximate general multivariate probability density functions by deterministic sample sets. For optimal sampling, the closeness to the given continuous density has to be assessed. This is a difficult challenge in multivariate…

Systems and Control · Electrical Eng. & Systems 2020-01-01 Uwe D. Hanebeck