English
Related papers

Related papers: Using Perturbation to Improve Goodness-of-Fit Test…

200 papers

We study sequences of scaled edge-corrected empirical (generalized) K-functions (modifying Ripley's K-function) each of them constructed from a single observation of a $d$-dimensional fourth-order stationary point process in a sampling…

Statistics Theory · Mathematics 2017-06-06 Lothar Heinrich

Nonparametric tests via kernel embedding of distributions have witnessed a great deal of practical successes in recent years. However, statistical properties of these tests are largely unknown beyond consistency against a fixed alternative.…

Statistics Theory · Mathematics 2019-09-10 Tong Li , Ming Yuan

Huge scale machine learning problems are nowadays tackled by distributed optimization algorithms, i.e. algorithms that leverage the compute power of many devices for training. The communication overhead is a key bottleneck that hinders…

Machine Learning · Computer Science 2018-11-30 Sebastian U. Stich , Jean-Baptiste Cordonnier , Martin Jaggi

In many biological applications, the primary objective of study is to quantify the magnitude of treatment effect between two groups. Cohens'd or strictly standardized mean difference (SSMD) can be used to measure effect size however, it is…

Applications · Statistics 2020-11-18 Seongyong Park , Shujaat Khan , Muhammad Moinuddin , Ubaid M. Al-Saggaf

We characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test…

Information Theory · Computer Science 2021-02-08 Shengyu Zhu , Biao Chen , Zhitang Chen , Pengfei Yang

Particle-based approximate Bayesian inference approaches such as Stein Variational Gradient Descent (SVGD) combine the flexibility and convergence guarantees of sampling methods with the computational benefits of variational inference. In…

Machine Learning · Computer Science 2021-07-30 Lauro Langosco di Langosco , Vincent Fortuin , Heiko Strathmann

Markov chain Monte Carlo samplers produce dependent streams of variates drawn from the limiting distribution of the Markov chain. With this as motivation, we introduce novel univariate kernel density estimators which are appropriate for the…

Methodology · Statistics 2016-07-29 Hang J. Kim , Steven N. MacEachern , Yoonsuh Jung

Recent years have seen a surge in methods for two-sample testing, among which the Maximum Mean Discrepancy (MMD) test has emerged as an effective tool for handling complex and high-dimensional data. Despite its success and widespread…

Machine Learning · Statistics 2026-05-21 Ikjun Choi , Ilmun Kim

We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, maximum mean…

Methodology · Statistics 2013-11-13 Dino Sejdinovic , Bharath Sriperumbudur , Arthur Gretton , Kenji Fukumizu

In this paper, we bound the error induced by using a weighted skeletonization of two data sets for computing a two sample test with kernel maximum mean discrepancy. The error is quantified in terms of the speed in which heat diffuses from…

Machine Learning · Statistics 2018-12-12 Alexander Cloninger

Representing, comparing, and measuring the distance between probability distributions is a key task in computational statistics and machine learning. The choice of representation and the associated distance determine properties of the…

Machine Learning · Statistics 2026-02-26 Masha Naslidnyk

Diffusion models are widely used as priors in imaging inverse problems. However, their performance often degrades under distribution shifts between the training and test-time images. Existing methods for identifying and quantifying…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Shirin Shoushtari , Edward P. Chandler , Yuanhao Wang , M. Salman Asif , Ulugbek S. Kamilov

One of the major problems in Machine Learning (ML) and Artificial Intelligence (AI) is the fact that the probability distribution of the test data in the real world could deviate substantially from the probability distribution of the…

Machine Learning · Computer Science 2025-10-21 Ozan K. Tonguz , Federico Taschin

Studying the stability of partially observed Markov decision processes (POMDPs) with respect to perturbations in either transition or observation kernels is a significant problem. While asymptotic robustness/stability results as approximate…

Optimization and Control · Mathematics 2025-09-15 Yunus Emre Demirci , Ali Devran Kara , Serdar Yüksel

We consider training and testing on mixture distributions with different training and test proportions. We show that in many settings, and in some sense generically, distribution shift can be beneficial, and test performance can improve due…

Machine Learning · Computer Science 2025-11-11 Marko Medvedev , Kaifeng Lyu , Zhiyuan Li , Nathan Srebro

This paper derives asymptotic approximations to the power of Cramer-von Mises (CvM) style tests for inference on a finite dimensional parameter defined by conditional moment inequalities in the case where the parameter is set identified.…

Applications · Statistics 2017-07-10 Timothy B. Armstrong

In this study, we establish a basis for selecting similarity measures when applying machine learning techniques to solve materials science problems. This selection is considered with an emphasis on the distinctiveness between materials that…

Machine Learning · Computer Science 2019-03-27 Tran-Thai Dang , Tien-Lam Pham , Hiori Kino , Takashi Miyake , Hieu-Chi Dam

In real supervised learning scenarios, it is not uncommon that the training and test sample follow different probability distributions, thus rendering the necessity to correct the sampling bias. Focusing on a particular covariate shift…

Machine Learning · Computer Science 2012-06-22 Yaoliang Yu , Csaba Szepesvari

Kernel means are frequently used to represent probability distributions in machine learning problems. In particular, the well known kernel density estimator and the kernel mean embedding both have the form of a kernel mean. Unfortunately,…

Machine Learning · Statistics 2015-03-03 E. Cruz Cortés , C. Scott

Fine-tuning pre-trained language models on downstream tasks with varying random seeds has been shown to be unstable, especially on small datasets. Many previous studies have investigated this instability and proposed methods to mitigate it.…

Computation and Language · Computer Science 2023-10-03 Yupei Du , Dong Nguyen