English
Related papers

Related papers: Two-sample KS test with approxQuantile in Apache S…

200 papers

The posterior predictive $p$-value (ppp) is widely used in Bayesian model evaluation. However, due to double use of the data, the ppp may not be a valid $p$-value even in large samples: The asymptotic null distribution of the ppp can be…

Statistics Theory · Mathematics 2026-01-13 Yueming Shen , Surya Tokdar

Over the last decade, an approach that has gained a lot of popularity to tackle nonparametric testing problems on general (i.e., non-Euclidean) domains is based on the notion of reproducing kernel Hilbert space (RKHS) embedding of…

Statistics Theory · Mathematics 2024-05-03 Omar Hagrass , Bharath K. Sriperumbudur , Bing Li

Nowadays, data analysis in the world of Big Data is connected typically to data mining, descriptive or exploratory statistics, e.~g.\ cluster analysis, classification or regression analysis. Aside these techniques there is a huge area of…

Applications · Statistics 2018-10-24 Taras Lazariv , Christoph Lehmann

This paper addresses the problem of approximating an unknown function from point evaluations. When obtaining these point evaluations is costly, minimising the required sample size becomes crucial, and it is unreasonable to reserve a…

Numerical Analysis · Mathematics 2025-11-06 Nando Hegemann , Anthony Nouy , Philipp Trunschke

We propose a nonparametric two-sample test procedure based on Maximum Mean Discrepancy (MMD) for testing the hypothesis that two samples of functions have the same underlying distribution, using kernels defined on function spaces. This…

Statistics Theory · Mathematics 2020-10-20 George Wynne , Andrew B. Duncan

We develop a kernel projected Wasserstein distance for the two-sample test, an essential building block in statistics and machine learning: given two sets of samples, to determine whether they are from the same distribution. This method…

Statistics Theory · Mathematics 2022-05-10 Jie Wang , Rui Gao , Yao Xie

In this paper we introduce a kernel-based measure for detecting differences between two conditional distributions. Using the `kernel trick' and nearest-neighbor graphs, we propose a consistent estimate of this measure which can be computed…

Methodology · Statistics 2024-08-30 Anirban Chatterjee , Ziang Niu , Bhaswar B. Bhattacharya

We consider conditional tests for non-negative discrete exponential families. We develop two Markov Chain Monte Carlo (MCMC) algorithms which allow us to sample from the conditional space and to perform approximated tests. The first…

Computation · Statistics 2017-07-27 Roberto Fontana , Francesca Romana Crucinio

The fidelity of financial market simulation is restricted by the so-called "non-identifiability" difficulty when calibrating high-frequency data. This paper first analyzes the inherent loss of data information in this difficulty, and…

Computational Engineering, Finance, and Science · Computer Science 2025-04-02 Peng Yang , Junji Ren , Feng Wang , Ke Tang

We characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test…

Information Theory · Computer Science 2021-02-08 Shengyu Zhu , Biao Chen , Zhitang Chen , Pengfei Yang

We propose a class of kernel-based two-sample tests, which aim to determine whether two sets of samples are drawn from the same distribution. Our tests are constructed from kernels parameterized by deep neural nets, trained to maximize test…

Machine Learning · Statistics 2021-01-15 Feng Liu , Wenkai Xu , Jie Lu , Guangquan Zhang , Arthur Gretton , Danica J. Sutherland

We analyze in detail the two-dimensional Kolmogorov-Smirnov test as a tool to learn about the distribution of the sources of the ultra-high energy cosmic rays. We confront in particular models based on AGN observed in X rays, on galaxies…

Astrophysics · Physics 2009-11-13 Diego Harari , Silvia Mollerach , Esteban Roulet

Approximate Markov chain Monte Carlo (MCMC) offers the promise of more rapid sampling at the cost of more biased inference. Since standard MCMC diagnostics fail to detect these biases, researchers have developed computable Stein discrepancy…

Machine Learning · Statistics 2020-10-16 Jackson Gorham , Lester Mackey

The advent of high dimensional single cell data in the biomedical sciences has necessitated the development of dimensionality-reduction tools. t-SNE and UMAP are the two most frequently used approaches, allowing clear visualisation of…

Quantitative Methods · Quantitative Biology 2021-12-09 Carlos P. Roca1 , Oliver T. Burton , Julika Neumann , Samar Tareen , Carly E. Whyte , Stéphanie Humblet-Baron , Adrian Liston

In the context of the widely used competing risks set-up we discuss different inference procedures for testing equality of two cumulative incidence functions, where the data may be subject to independent right-censoring or left-truncation.…

Statistics Theory · Mathematics 2015-10-13 Dennis Dobler , Markus Pauly

Given an i.i.d. sample drawn from a density $f$, we propose to test that $f$ equals some prescribed density $f_0$ or that $f$ belongs to some translation/scale family. We introduce a multiple testing procedure based on an estimation of the…

Statistics Theory · Mathematics 2016-08-16 Magalie Fromont , Béatrice Laurent

A common disadvantage in existing distribution-free two-sample testing approaches is that the computational complexity could be high. Specifically, if the sample size is $N$, the computational complexity of those two-sample tests is at…

Methodology · Statistics 2017-07-18 Cheng Huang , Xiaoming Huo

Two-sample testing is a fundamental problem in statistics. Despite its long history, there has been renewed interest in this problem with the advent of high-dimensional and complex data. Specifically, in the machine learning literature,…

Methodology · Statistics 2019-11-19 Ilmun Kim , Ann B. Lee , Jing Lei

The two-sample hypothesis testing problem is studied for the challenging scenario of high dimensional data sets with small sample sizes. We show that the two-sample hypothesis testing problem can be posed as a one-class set classification…

Machine Learning · Statistics 2017-11-15 Hamed Masnadi-Shirazi

Kernel two-sample tests have been widely used, and the development of efficient methods for high-dimensional, large-scale data is receiving increasing attention in the big data era. However, existing methods, such as the maximum mean…

Methodology · Statistics 2025-10-03 Hoseung Song , Hao Chen