English
Related papers

Related papers: On High Dimensional Behaviour of Some Two-Sample T…

200 papers

Consider $d$ dependent change point tests, each based on a CUSUM-statistic. We provide an asymptotic theory that allows us to deal with the maximum over all test statistics as both the sample size $n$ and $d$ tend to infinity. We achieve…

Statistics Theory · Mathematics 2017-12-07 Moritz Jirak

For the mean vector test in high dimension, Ayyala et al.(2017,153:136-155) proposed new test statistics when the observational vectors are M dependent. Under certain conditions, the test statistics for one-same and two-sample cases were…

Statistics Theory · Mathematics 2019-04-23 Seonghun Cho , Johan Lim , Deepak Nag Ayyala , Junyong Park , Anindya Roy

Classification and clustering are both important topics in statistical learning. A natural question herein is whether predefined classes are really different from one another, or whether clusters are really there. Specifically, we may be…

Machine Learning · Statistics 2015-09-22 Qiyi Lu , Xingye Qiao

We study two-sample tests for relevant differences in persistence diagrams obtained from $L^p$-$m$-approximable data $(\mathcal{X}_t)_t$ and $(\mathcal{Y}_t)_t$. To this end, we compare variance estimates w.r.t.\ the Wasserstein metrics on…

Statistics Theory · Mathematics 2024-01-22 Johannes Krebs , Daniel Rademacher

In this paper, we study the problem of testing the mean vectors of high dimensional data in both one-sample and two-sample cases. The proposed testing procedures employ maximum-type statistics and the parametric bootstrap techniques to…

Statistics Theory · Mathematics 2018-01-23 Jinyuan Chang , Chao Zheng , Wen-Xin Zhou , Wen Zhou

Maximum Variance Unfolding is one of the main methods for (nonlinear) dimensionality reduction. We study its large sample limit, providing specific rates of convergence under standard assumptions. We find that it is consistent when the…

Machine Learning · Statistics 2013-03-21 Ery Arias-Castro , Bruno Pelletier

The aim of this article is to prove strong convergence results on the difference between the solution to highly oscillatory problems posed in thin domains and its two-scale expansion. We first consider the case of the linear diffusion…

Analysis of PDEs · Mathematics 2025-07-29 Virginie Ehrlacher , Arthur Lebée , Frédéric Legoll , Adrien Lesage

Classical asymptotic theory for statistical inference usually involves calibrating a statistic by fixing the dimension $d$ while letting the sample size $n$ increase to infinity. Recently, much effort has been dedicated towards…

Statistics Theory · Mathematics 2024-05-14 Ilmun Kim , Aaditya Ramdas

Two-sample tests are important areas aiming to determine whether two collections of observations follow the same distribution or not. We propose two-sample tests based on integral probability metric (IPM) for high-dimensional samples…

Machine Learning · Statistics 2023-04-21 Jie Wang , Minshuo Chen , Tuo Zhao , Wenjing Liao , Yao Xie

We consider an analysis of variance type problem, where the sample observations are random elements in an infinite dimensional space. This scenario covers the case, where the observations are random functions. For such a problem, we propose…

Methodology · Statistics 2022-07-26 Joydeep Chowdhury , Probal Chaudhuri

Over the last decade, an approach that has gained a lot of popularity to tackle nonparametric testing problems on general (i.e., non-Euclidean) domains is based on the notion of reproducing kernel Hilbert space (RKHS) embedding of…

Statistics Theory · Mathematics 2024-05-03 Omar Hagrass , Bharath K. Sriperumbudur , Bing Li

Central limit theorems (CLTs) for high-dimensional random vectors with dimension possibly growing with the sample size have received a lot of attention in the recent times. Chernozhukov et al. (2017) proved a Berry--Esseen type result for…

Statistics Theory · Mathematics 2019-06-26 Arun Kumar Kuchibhotla , Somabha Mukherjee , Debapratim Banerjee

Methods of performing anomaly detection on high-dimensional data sets are needed, since algorithms which are trained on data are only expected to perform well on data that is similar to the training data. There are theoretical results on…

Machine Learning · Computer Science 2020-11-13 Forrest Laine , Claire Tomlin

In this paper we propose a new test of heteroscedasticity for parametric regression models and partial linear regression models in high dimensional settings. When the dimension of covariates is large, existing tests of heteroscedasticity…

Methodology · Statistics 2018-08-09 Falong Tan , Xuejun Jiang , Xu Guo , Lixing Zhu

Change-point detection has been a classical problem in statistics and econometrics. This work focuses on the problem of detecting abrupt distributional changes in the data-generating distribution of a sequence of high-dimensional…

Methodology · Statistics 2021-05-20 Shubhadeep Chakraborty , Xianyang Zhang

We present the results of a large number of simulation studies regarding the power of various goodness-of-fit as well as non-parametric two-sample tests for multivariate data. In two dimensions this includes both continuous and discrete…

Methodology · Statistics 2026-05-13 Wolfgang Rolke

A non parametric method based on the empirical likelihood is proposed for detecting the change in the coefficients of high-dimensional linear model where the number of model variables may increase as the sample size increases. This amounts…

Statistics Theory · Mathematics 2015-06-22 Gabriela Ciuperca , Zahraa Salloum

We present a simple method for assessing the predictive performance of high-dimensional models directly in data space when only samples are available. Our approach is to compare the quantiles of observables predicted by a model to those of…

Instrumentation and Methods for Astrophysics · Physics 2025-01-16 Stephen Thorp , Hiranya V. Peiris , Daniel J. Mortlock , Justin Alsing , Boris Leistedt , Sinan Deger

Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more…

Machine Learning · Statistics 2025-04-25 Iskander Azangulov , George Deligiannidis , Judith Rousseau

We use a suitable version of the so-called "kernel trick" to devise two-sample (homogeneity) tests, especially focussed on high-dimensional and functional data. Our proposal entails a simplification related to the important practical…

Statistics Theory · Mathematics 2024-04-24 Javier Cárcamo , Antonio Cuevas , Luis-Alberto Rodríguez