Related papers: Robust estimations from distribution structures: I…
We provide a unified framework for independence and mean independence tests based on the Hilbert-Schmidt independence criterion, extending some previous results in the literature to hold in general topological spaces. We also present a…
Covariate shift occurs prevalently in practice, where the input distributions of the source and target data are substantially different. Despite its practical importance in various learning problems, most of the existing methods only focus…
Given $M \geq 2$ distributions defined on a general measurable space, we introduce a nonparametric (kernel) measure of multi-sample dissimilarity (KMD) -- a parameter that quantifies the difference between the $M$ distributions. The…
We develop Gaussian approximations for high-dimensional vectors formed by second-order $U$- and $V$-statistics whose kernels depend on sample size under independent but not identically distributed (i.n.i.d.) sampling. Our results hold…
This work presents the concept of kernel mean embedding and kernel probabilistic programming in the context of stochastic systems. We propose formulations to represent, compare, and propagate uncertainties for fairly general stochastic…
We present a robust test for change-points in time series which is based on the two-sample Hodges-Lehmann estimator. We develop new limit theory for a class of statistics based on the two-sample U-quantile processes, in the case of short…
We establish the limiting spectral distribution of Kendall's correlation matrices in the moderate high-dimensional regime where the dimension grows slower than the sample size. Our framework allows observations to be independent but not…
Estimating the score, i.e., the gradient of log density function, from a set of samples generated by an unknown distribution is a fundamental task in inference and learning of probabilistic models that involve flexible yet intractable…
These notes provide a self-contained introduction to kernel methods and their geometric foundations in machine learning. Starting from the construction of Hilbert spaces, we develop the theory of positive definite kernels, reproducing…
Seismic noise cross correlations are used to image crustal structure and heterogeneity. Typically, seismic networks are only anisotropically illuminated by seismic noise, a consequence of the non-uniform distribution of sources. Here, we…
The proportion of false null hypotheses is a very important quantity in statistical modelling and inference based on the two-component mixture model and its extensions, and in control and estimation of the false discovery rate and false…
Are two sets of observations drawn from the same distribution? This problem is a two-sample test. Kernel methods lead to many appealing properties. Indeed state-of-the-art approaches use the $L^2$ distance between kernel-based distribution…
We study the problem of testing, using only a single sample, between mean field distributions (like Curie-Weiss, Erd\H{o}s-R\'enyi) and structured Gibbs distributions (like Ising model on sparse graphs and Exponential Random Graphs). Our…
Motivated by small bandwidth asymptotics for kernel-based semiparametric estimators in econometrics, this paper establishes Gaussian approximation results for high-dimensional fixed-order $U$-statistics whose kernels depend on the sample…
Kernel mean embeddings, a widely used technique in machine learning, map probability distributions to elements of a reproducing kernel Hilbert space (RKHS). For supervised learning problems, where input-output pairs are observed, the…
The global clustering coefficient is an effective measure for analyzing and comparing the structures of complex networks. The random annulus graph is a modified version of the well-known Erd\H{o}s-R\'{e}nyi random graph. It has been…
We trace a conceptual genealogy from Abraham de Moivre's derivation of the normal curve (1733) to the modern distributional approach to statistics. De Moivre's Approximatio ad Summam Terminorum Binomii gave the first systematic derivation…
We propose a data-driven approach to quantify the uncertainty of models constructed by kernel methods. Our approach minimizes the needed distributional assumptions, hence, instead of working with, for example, Gaussian processes or…
The notion of maximal-spacing in several dimensions was introduced and studied by Deheuvels (1983) for data uniformly distributed on the unit cube. Later on, Janson (1987) extended the results to data uniformly distributed on any bounded…
We consider the problem of learning a set from random samples. We show how relevant geometric and topological properties of a set can be studied analytically using concepts from the theory of reproducing kernel Hilbert spaces. A new kind of…