English
Related papers

Related papers: Accelerated Computation of a High Dimensional Kolm…

200 papers

The Kolmogorov-Smirnov (KS) test is a nonparametric statistical test used to test for differences between univariate probability distributions. The versatility of the KS test has made it a cornerstone of statistical analysis across many…

Methodology · Statistics 2022-11-21 Connor Puritz , Elan Ness-Cohn , Rosemary Braun

We revisit extending the Kolmogorov-Smirnov distance between probability distributions to the multidimensional setting and make new arguments about the proper way to approach this generalization. Our proposed formulation maximizes the…

Computation · Statistics 2025-04-16 Peter Matthew Jacobs , Foad Namjoo , Jeff M. Phillips

Kernel two-sample tests have been widely used, and the development of efficient methods for high-dimensional, large-scale data is receiving increasing attention in the big data era. However, existing methods, such as the maximum mean…

Methodology · Statistics 2025-10-03 Hoseung Song , Hao Chen

The two-sample Kolmogorov-Smirnov test is a widely used statistical test for detecting whether two samples are likely to come from the same distribution. Implementations typically recur on an article of Hodges from 1957. The advances in…

Computation · Statistics 2021-09-27 Thomas Viehmann

Big Data has become an ever more commonplace setting that is encountered by data analysts. In the Big Data setting, analysts are faced with very large numbers of observations as well as data that arrive as a stream, both of which are…

Computation · Statistics 2017-04-13 Hien Duy Nguyen

The classical two-sample test of Kolmogorov-Smirnov (KS) is widely used to test whether empirical samples come from the same distribution. Even though most statistical packages provide an implementation, carrying out the test in big data…

Computation · Statistics 2023-12-18 Bradley Eck , Duygu Kabakci-Zorlu , Amadou Ba

One of the major problems in Machine Learning (ML) and Artificial Intelligence (AI) is the fact that the probability distribution of the test data in the real world could deviate substantially from the probability distribution of the…

Machine Learning · Computer Science 2025-10-21 Ozan K. Tonguz , Federico Taschin

The boom of DL technology leads to massive DL models built and shared, which facilitates the acquisition and reuse of DL models. For a given task, we encounter multiple DL models available with the same functionality, which are considered…

Software Engineering · Computer Science 2021-03-10 Linghan Meng , Yanhui Li , Lin Chen , Zhi Wang , Di Wu , Yuming Zhou , Baowen Xu

This paper proposes using a method named Double Score Matching (DSM) to do mass-imputation and presents an application to make inferences with a nonprobability sample. DSM is a $k$-Nearest Neighbors algorithm that uses two balance scores…

Methodology · Statistics 2021-10-19 Ali Furkan Kalay

In many modern data sets, High dimension low sample size (HDLSS) data is prevalent in many fields of studies. There has been an increased focus recently on using machine learning and statistical methods to mine valuable information out of…

Optimization and Control · Mathematics 2023-05-23 Srivathsan Amruth , Xin Yee Lam

We present an extension of the Kolmogorov-Smirnov (KS) two-sample test, which can be more sensitive to differences in the tails. Our test statistic is an integral probability metric (IPM) defined over a higher-order total variation ball,…

Machine Learning · Statistics 2019-03-26 Veeranjaneyulu Sadhanala , Yu-Xiang Wang , Aaditya Ramdas , Ryan J. Tibshirani

Testing for the equality of two high-dimensional distributions is a challenging problem, and this becomes even more challenging when the sample size is small. Over the last few decades, several graph-based two-sample tests have been…

Methodology · Statistics 2019-11-22 Soham Sarkar , Rahul Biswas , Anil K. Ghosh

We extend the Kolmogorov--Smirnov (K-S) test to multiple dimensions by suggesting a $\mathbb{R}^n \rightarrow [0,1]$ mapping based on the probability content of the highest probability density region of the reference distribution under…

Instrumentation and Methods for Astrophysics · Physics 2015-05-18 Diana Harrison , David Sutton , Pedro Carvalho , Michael Hobson

Kernel two-sample tests have been widely used for multivariate data to test equality of distributions. However, existing tests based on mapping distributions into a reproducing kernel Hilbert space mainly target specific alternatives and do…

Methodology · Statistics 2023-11-21 Hoseung Song , Hao Chen

Two-sample hypothesis testing is a fundamental problem with various applications, which faces new challenges in the high-dimensional context. To mitigate the issue of the curse of dimensionality, high-dimensional data are typically assumed…

Methodology · Statistics 2026-04-06 Jiaqi Gu , Ruoxu Tan , Guosheng Yin

Statistical distances quantifies the difference between two statistical constructs. In this article, we describe reference values for a distance between samples derived from the Kolmogorov-Smirnov statistic $D_{F,F'}$. Each measure of the…

Data Analysis, Statistics and Probability · Physics 2017-11-03 Renato Fabbri , Fernando Gularte De León

Data representation techniques have made a substantial contribution to advancing data processing and machine learning (ML). Improving predictive power was the focus of previous representation techniques, which unfortunately perform rather…

Machine Learning · Computer Science 2022-05-24 Qiyou Duan , Hadi Ghauch , Taejoon Kim

In this article, we propose some two-sample tests based on ball divergence and investigate their high dimensional behavior. First, we study their behavior for High Dimension, Low Sample Size (HDLSS) data, and under appropriate regularity…

Statistics Theory · Mathematics 2024-10-08 Bilol Banerjee , Anil K. Ghosh

This paper investigates the estimation of the self-similarity parameter in fractional processes. We re-examine the Kolmogorov-Smirnov (KS) test as a distribution-based method for assessing self-similarity, emphasizing its robustness and…

Methodology · Statistics 2025-02-12 Daniele Angelini , Sergio Bianchi

In high-dimension, low-sample size (HDLSS) data, it is not always true that closeness of two objects reflects a hidden cluster structure. We point out the important fact that it is not the closeness, but the "values" of distance that…

Machine Learning · Statistics 2013-12-30 Yoshikazu Terada
‹ Prev 1 2 3 10 Next ›