English
Related papers

Related papers: A unified framework for multivariate two-sample an…

200 papers

In kernel methods, the median heuristic has been widely used as a way of setting the bandwidth of RBF kernels. While its empirical performances make it a safe choice under many circumstances, there is little theoretical understanding of why…

Statistics Theory · Mathematics 2018-10-31 Damien Garreau , Wittawat Jitkrittum , Motonobu Kanagawa

We consider marked empirical processes indexed by a randomly projected functional covariate to construct goodness-of-fit tests for the functional linear model with scalar response. The test statistics are built from continuous functionals…

Accurately specifying covariance structures is critical for valid inference in longitudinal and functional data analysis, particularly when data are sparsely observed. In this study, we develop a global goodness-of-fit test to assess…

Methodology · Statistics 2025-03-31 Dhrubajyoti Ghosh , Zhuolin Song , Luo Xiao , Sheng Luo

Spherical and hyperspherical data are commonly encountered in diverse applied research domains, underscoring the vital task of assessing independence within such data structures. In this context, we investigate the properties of test…

Methodology · Statistics 2024-01-23 Marija Cuparić , Bruno Ebner , Bojana Milošević

Regression tasks, notably in safety-critical domains, require proper uncertainty quantification, yet the literature remains largely classification-focused. In this light, we introduce a family of measures for total, aleatoric, and epistemic…

Machine Learning · Computer Science 2025-10-30 Christopher Bülte , Yusuf Sale , Gitta Kutyniok , Eyke Hüllermeier

A goodness-of-fit test for one-parameter count distributions with finite second moment is proposed. The test statistic is derived from the $L^1$ distance of a function of the probability generating function of the model under the null…

Statistics Theory · Mathematics 2024-06-11 Antonio Di Noia , Lucio Barabesi , Marzia Marcheselli , Caterina Pisani , Luca Pratelli

We develop a nonparametric two-sample test for distributions supported on the cone of symmetric positive definite matrices. The procedure relies on the Wishart kernel density estimator (KDE) introduced by Belzile et al. (2025), whose…

Statistics Theory · Mathematics 2026-03-17 Frédéric Ouimet

The Maximum Mean Discrepancy (MMD) has been the state-of-the-art nonparametric test for tackling the two-sample problem. Its statistic is given by the difference in expectations of the witness function, a real-valued function defined as a…

Machine Learning · Computer Science 2022-02-14 Jonas M. Kübler , Wittawat Jitkrittum , Bernhard Schölkopf , Krikamol Muandet

In this work, a goodness-of-fit test for the null hypothesis of a functional linear model with scalar response is proposed. The test is based on a generalization to the functional framework of a previous one, designed for the…

The two-sample hypothesis testing problem is studied for the challenging scenario of high dimensional data sets with small sample sizes. We show that the two-sample hypothesis testing problem can be posed as a one-class set classification…

Machine Learning · Statistics 2017-11-15 Hamed Masnadi-Shirazi

This paper develops a smooth test of goodness-of-fit for elliptical distributions. The test is adaptively omnibus, invariant to affine-linear transformations and has a convenient expression that can be broken into components. These…

Statistics Theory · Mathematics 2019-02-12 Gilles R. Ducharme , Pierre Lafaye de Micheaux

In this work, the distributional properties of the goodness-of-fit term in likelihood-based information criteria are explored. These properties are then leveraged to construct a novel goodness-of-fit test for normal linear regression models…

Methodology · Statistics 2023-09-20 Scott H. Koeneman , Joseph E. Cavanaugh

The Maximum Mean Discrepancy (MMD) is a widely used multivariate distance metric for two-sample testing. The standard MMD test statistic has an intractable null distribution typically requiring costly resampling or permutation approaches…

Methodology · Statistics 2026-02-24 Anirban Chatterjee , Aaditya Ramdas

Kernelized Stein discrepancy (KSD) is a score-based discrepancy widely used in goodness-of-fit tests. It can be applied even when the target distribution has an unknown normalising factor, such as in Bayesian analysis. We show theoretically…

Machine Learning · Statistics 2023-06-06 Xing Liu , Andrew B. Duncan , Axel Gandy

We characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test…

Information Theory · Computer Science 2021-02-08 Shengyu Zhu , Biao Chen , Zhitang Chen , Pengfei Yang

The Bayesian nonparametric inference and Dirichlet process are popular tools in statistical methodologies. In this paper, we employ the Dirichlet process in hypothesis testing to propose a Bayesian nonparametric chi-squared goodness-of-fit…

Statistics Theory · Mathematics 2016-06-20 Reyhaneh Hosseini , Mahmoud Zarepour

A distance measure is presented between two unitary propagators of quantum systems of differing dimensions along with a corresponding method of computation. A typical application is to compare the propagator of the actual (real) process…

Quantum Physics · Physics 2007-05-23 Robert L. Kosut , Matthew Grace , Constantin Brif , Herschel Rabitz

We develop here several goodness-of-fit tests for testing the k-monotonicity of a discrete density, based on the empirical distribution of the observations. Our tests are non-parametric, easy to implement and are proved to be asymptotically…

Methodology · Statistics 2017-08-30 Jade Giguelay , Sylvie Huet

Despite a substantial literature on nonparametric two-sample goodness-of-fit testing in arbitrary dimensions spanning decades, there is no mention there of any curse of dimensionality. Only more recently Ramdas et al. (2015) have discussed…

Statistics Theory · Mathematics 2018-09-13 Ery Arias-Castro , Bruno Pelletier , Venkatesh Saligrama

Distance-based clustering and classification are widely used in various fields to group mixed numeric and categorical data. In many algorithms, a predefined distance measurement is used to cluster data points based on their dissimilarity.…

Machine Learning · Computer Science 2024-10-14 Jesse S. Ghashti , John R. J. Thompson