English
Related papers

Related papers: A Permutation-free Kernel Two-Sample Test

200 papers

As a novel similarity measure that is defined as the expectation of a kernel function between two random variables, correntropy has been successfully applied in robust machine learning and signal processing to combat large outliers. The…

Machine Learning · Computer Science 2021-09-07 Badong Chen , Yuqing Xie , Xin Wang , Zejian yuan , Pengju Ren , Jing Qin

Conditional independence testing is a fundamental problem underlying causal discovery and a particularly challenging task in the presence of nonlinear and high-dimensional dependencies. Here a fully non-parametric test for continuous data…

Machine Learning · Statistics 2017-09-06 Jakob Runge

We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However,…

Machine Learning · Statistics 2026-05-05 Gyumin Lee , Shubhanshu Shekhar , Ilmun Kim

The mean shift (MS) is a non-parametric, density-based, iterative algorithm with prominent usage in clustering and image segmentation. A rigorous proof for the convergence of its mode estimate sequence in full generality remains unknown. In…

Machine Learning · Statistics 2026-03-17 Susovan Pal

The modelling of data on a spherical surface requires the consideration of directional probability distributions. To model asymmetrically distributed data on a three-dimensional sphere, Kent distributions are often used. The moment…

Machine Learning · Computer Science 2015-06-29 Parthan Kasarapu

The measurement-device-independent quantum key distribution (MDI-QKD) can be immune to all detector side-channel attacks. Moreover, it can be easily implemented combining with the matured decoy-state methods under current technology. It…

Quantum Physics · Physics 2021-08-11 Yi-Peng Chen , Jing-Yang Liu , Mingshuo Sun , Xing-Xu Zhou , Chun-Hui Zhang , Jian Li , Qin Wang

Many interesting machine learning problems are best posed by considering instances that are distributions, or sample sets drawn from distributions. Previous work devoted to machine learning tasks with distributional inputs has done so…

Machine Learning · Statistics 2021-01-15 Danica J. Sutherland , Junier B. Oliva , Barnabás Póczos , Jeff Schneider

In this article, we present a nonparametric method for the general two-sample problem involving functional random variables modelled as elements of a separable Hilbert space ${\cal H}$. First, we present a general recipe based on linear…

Methodology · Statistics 2024-10-08 Bilol Banerjee

Kernel mean embeddings have recently attracted the attention of the machine learning community. They map measures $\mu$ from some set $M$ to functions in a reproducing kernel Hilbert space (RKHS) with kernel $k$. The RKHS distance of two…

Machine Learning · Statistics 2019-12-18 Carl-Johann Simon-Gabriel , Bernhard Schölkopf

We study the problem of quickest detection of a change in the mean of an observation sequence, under the assumption that both the pre- and post-change distributions have bounded support. We first study the case where the pre-change…

Signal Processing · Electrical Eng. & Systems 2021-01-15 Yuchen Liang , Venugopal V. Veeravalli

In this paper, we bound the error induced by using a weighted skeletonization of two data sets for computing a two sample test with kernel maximum mean discrepancy. The error is quantified in terms of the speed in which heat diffuses from…

Machine Learning · Statistics 2018-12-12 Alexander Cloninger

This paper investigates the utilization of maximum and average distance correlations for multivariate independence testing. We characterize their consistency properties in high-dimensional settings with respect to the number of marginally…

Machine Learning · Statistics 2025-06-11 Cencheng Shen , Yuexiao Dong

This paper considers the change point detection problem under dependent samples. In particular, we provide performance guarantees for the MMD-CUSUM test under exponentially $\alpha$, $\beta$, and fast $\phi$-mixing processes, which…

Systems and Control · Electrical Eng. & Systems 2024-05-13 Hao Chen , Abhishek Gupta , Yin Sun , Ness Shroff

Exploring the free-energy landscape along reaction coordinates or system parameters $\lambda$ is central to many studies of high-dimensional model systems in physics, e.g. large molecules or spin glasses. In simulations this usually…

Statistical Mechanics · Physics 2018-09-05 Viveca Lindahl , Jack Lidmar , Berk Hess

Local polynomial density (LPD) estimators are widely used for inference on boundary features of the density function. Contrary to conventional wisdom, we show that kernel choice substantially affects efficiency. Theory, simulations, and…

Econometrics · Economics 2026-01-08 Shunsuke Imai , Yuta Okamoto

We revisit extending the Kolmogorov-Smirnov distance between probability distributions to the multidimensional setting and make new arguments about the proper way to approach this generalization. Our proposed formulation maximizes the…

Computation · Statistics 2025-04-16 Peter Matthew Jacobs , Foad Namjoo , Jeff M. Phillips

To adapt kernel two-sample and independence testing to complex structured data, aggregation of multiple kernels is frequently employed to boost testing power compared to single-kernel tests. However, we observe a phenomenon that directly…

Machine Learning · Computer Science 2025-10-14 Zhijian Zhou , Xunye Tian , Liuhua Peng , Chao Lei , Antonin Schrab , Danica J. Sutherland , Feng Liu

We study query time bounds for the fundamental problem of estimating the kernel mean $\frac1{|X|}\sum_{x\in X}\mathbf{k}(x,y)$ of a query $y$ in a finite dataset $X\subset\mathbb{R}^d$ up to a prescribed additive error $\varepsilon$. The…

Data Structures and Algorithms · Computer Science 2026-05-05 Tal Wagner

Models like support vector machines or Gaussian process regression often require positive semi-definite kernels. These kernels may be based on distance functions. While definiteness is proven for common distances and kernels, a proof for a…

Machine Learning · Computer Science 2018-07-11 Martin Zaefferer , Thomas Bartz-Beielstein , Günter Rudolph

The standardized mean difference (SMD) is a widely used measure of effect size, particularly common in psychology, clinical trials, and meta-analysis involving continuous outcomes. Traditionally, under the equal variance assumption, the SMD…

Methodology · Statistics 2025-06-05 Jiandong Shi , Xiaochen Zhang , Lu Lin , Hiu Yee Kwan , Tiejun Tong
‹ Prev 1 8 9 10 Next ›