English
Related papers

Related papers: Two Sample Testing in High Dimension via Maximum M…

200 papers

This paper adresses the problem of testing for the equality of $k$ probability distributions on Hilbert spaces, with $k\geqslant 2$. We introduce a generalization of the maximum variance discrepancy called multiple maximum variance…

Statistics Theory · Mathematics 2024-04-16 Armando Sosthène Kali Balogoun , Guy Martial Nkiet

In repeated Measure Designs with multiple groups, the primary purpose is to compare different groups in various aspects. For several reasons, the number of measurements and therefore the dimension of the observation vectors can depend on…

Statistics Theory · Mathematics 2022-07-20 Paavo Sattler , Markus Pauly

Combinatorial optimization is a fertile testing ground for statistical physics methods developed in the context of disordered systems, allowing one to confront theoretical mean field predictions with actual properties of finite dimensional…

Disordered Systems and Neural Networks · Physics 2009-10-31 J. Houdayer , J. H. Boutet de Monvel , O. C. Martin

Classical asymptotic theory for statistical inference usually involves calibrating a statistic by fixing the dimension $d$ while letting the sample size $n$ increase to infinity. Recently, much effort has been dedicated towards…

Statistics Theory · Mathematics 2024-05-14 Ilmun Kim , Aaditya Ramdas

The stochastic block model is a popular tool for detecting community structures in network data. Detecting the difference between two community structures is an important issue for stochastic block models. However, the two-sample test has…

Methodology · Statistics 2022-12-21 Kang Fu , Jianwei Hu , Seydou Keita , Hao Liu

In this paper, we study inference for high-dimensional data characterized by small sample sizes relative to the dimension of the data. In particular, we provide an infinite-dimensional framework to study statistical models that involve…

Statistics Theory · Mathematics 2010-02-25 Jim Kuelbs , Anand N. Vidyashankar

The design of a metric between probability distributions is a longstanding problem motivated by numerous applications in Machine Learning. Focusing on continuous probability distributions on the Euclidean space $\mathbb{R}^d$, we introduce…

We give explicit theoretical and heuristical bounds for how big does a data set sampled from a reach-1 submanifold M of euclidian space need to be, to be able to estimate the dimension of M with 90% confidence.

Computational Geometry · Computer Science 2022-09-08 Lucien Grillet , Juan Souto

For a given metric measure space $(X,d,\mu)$ we consider finite samples of points, calculate the matrix of distances between them and then reconstruct the points in some finite-dimensional space using the multidimensional scaling (MDS)…

Metric Geometry · Mathematics 2022-08-02 Alexey Kroshnin , Eugene Stepanov , Dario Trevisan

We investigate unsupervised anomaly detection for high-dimensional data and introduce a deep metric learning (DML) based framework. In particular, we learn a distance metric through a deep neural network. Through this metric, we project the…

Machine Learning · Computer Science 2020-05-13 Selim F. Yilmaz , Suleyman S. Kozat

This paper is devoted to the statistical and numerical properties of the geometric median, and its applications to the problem of robust mean estimation via the median of means principle. Our main theoretical results include (a) an upper…

Statistics Theory · Mathematics 2023-07-21 Stanislav Minsker , Nate Strawn

Two-sample tests are important areas aiming to determine whether two collections of observations follow the same distribution or not. We propose two-sample tests based on integral probability metric (IPM) for high-dimensional samples…

Machine Learning · Statistics 2023-04-21 Jie Wang , Minshuo Chen , Tuo Zhao , Wenjing Liao , Yao Xie

Approximate Markov chain Monte Carlo (MCMC) offers the promise of more rapid sampling at the cost of more biased inference. Since standard MCMC diagnostics fail to detect these biases, researchers have developed computable Stein discrepancy…

Machine Learning · Statistics 2020-10-16 Jackson Gorham , Lester Mackey

We propose a new probabilistic characterization of the uniform distribution on the hypersphere in terms of the distribution of pairwise inner products, extending the ideas of \citep{cuesta2009projection,cuesta2007sharp} in a data-driven…

Statistics Theory · Mathematics 2026-04-14 Tiefeng Jiang , Tuan Pham

Evaluating generative adversarial networks (GANs) is inherently challenging. In this paper, we revisit several representative sample-based evaluation metrics for GANs, and address the problem of how to evaluate the evaluation metrics. We…

Machine Learning · Computer Science 2018-08-20 Qiantong Xu , Gao Huang , Yang Yuan , Chuan Guo , Yu Sun , Felix Wu , Kilian Weinberger

It has been a long history in testing whether a mean vector with a fixed dimension has a specified value. Some well-known tests include the Hotelling $T^2$-test and the empirical likelihood ratio test proposed by Owen [Biometrika 75 (1988)…

Methodology · Statistics 2014-05-21 Liang Peng , Yongcheng Qi , Fang Wang

Tests based on sample mean vectors and sample spatial signs have been studied in the recent literature for high dimensional data with the dimension larger than the sample size. For suitable sequences of alternatives, we show that the powers…

Statistics Theory · Mathematics 2015-05-22 Anirvan Chakraborty , Probal Chaudhuri

It is of great interest to test the equality of the means in two samples of functional data. Past research has predominantly concentrated on low-dimensional functional data, a focus that may not hold up in high-dimensional scenarios. In…

Methodology · Statistics 2025-05-27 Shouxia Wang , Jiguo Cao , Hua Liu , Jinhong You , Jicai Liu

Distance queries are a basic tool in data analysis. They are used for detection and localization of change for the purpose of anomaly detection, monitoring, or planning. Distance queries are particularly useful when data sets such as…

Data Structures and Algorithms · Computer Science 2015-03-20 Edith Cohen

This paper introduces an approach for detecting differences in the first-order structures of spatial point patterns. The proposed approach leverages the kernel mean embedding in a novel way by introducing its approximate version tailored to…

Methodology · Statistics 2020-06-15 Raif M. Rustamov , James T. Klosowski