Related papers: How many moments does MMD compare?
The Word Movers Distance (WMD) measures the semantic dissimilarity between two text documents by computing the cost of optimally moving all words of a source/query document to the most similar words of a target document. Computing WMD…
This paper introduces kdiff, a novel kernel-based measure for estimating distances between instances of time series, random fields and other forms of structured data. This measure is based on the idea of matching distributions that only…
The Maximum Mean Discrepancy (MMD) is a cornerstone statistic for nonparametric two-sample testing, but its test power is dictated entirely by the chosen kernel. Because any fixed kernel inherently fails to distinguish certain…
Approximate Markov chain Monte Carlo (MCMC) offers the promise of more rapid sampling at the cost of more biased inference. Since standard MCMC diagnostics fail to detect these biases, researchers have developed computable Stein discrepancy…
Kernel mean embeddings have recently attracted the attention of the machine learning community. They map measures $\mu$ from some set $M$ to functions in a reproducing kernel Hilbert space (RKHS) with kernel $k$. The RKHS distance of two…
Kernel embeddings of distributions and the Maximum Mean Discrepancy (MMD), the resulting distance between distributions, are useful tools for fully nonparametric two-sample testing and learning on distributions. However, it is rarely that…
We propose a nonparametric two-sample test procedure based on Maximum Mean Discrepancy (MMD) for testing the hypothesis that two samples of functions have the same underlying distribution, using kernels defined on function spaces. This…
Kernel techniques are among the most popular and flexible approaches in data science allowing to represent probability measures without loss of information under mild conditions. The resulting mapping called mean embedding gives rise to a…
The approximation of a discrete probability distribution $\mathbf{t}$ by an $M$-type distribution $\mathbf{p}$ is considered. The approximation error is measured by the informational divergence $\mathbb{D}(\mathbf{t}\Vert\mathbf{p})$, which…
Biclustering algorithms partition data and covariates simultaneously, providing new insights in several domains, such as analyzing gene expression to discover new biological functions. This paper develops a new model-free biclustering…
We characterize the reproducing kernel Hilbert spaces whose elements are $p$-integrable functions in terms of the boundedness of the integral operator whose kernel is the reproducing kernel. Moreover, for $p=2$ we show that the spectral…
The Maximum Mean Discrepancy (MMD) has been the state-of-the-art nonparametric test for tackling the two-sample problem. Its statistic is given by the difference in expectations of the witness function, a real-valued function defined as a…
Let $\{X_n\}_{n\in\N}$ be a Markov chain on a measurable space $\X$ with transition kernel $P$ and let $V:\X\r[1,+\infty)$. The Markov kernel $P$ is here considered as a linear bounded operator on the weighted-supremum space $\cB_V$…
In this article, we begin a systematic study of the boundedness and the nuclearity properties of multilinear periodic pseudo-differential operators and multilinear discrete pseudo-differential operators on $L^p$-spaces. First, we prove…
We study a new class of pseudo differential operators whose symbols satisfy the differential inequality with a mixture of homogeneities. On the other hand, by taking singular integral realization, it can be equivalently defined by kernels…
We deal with kernel theorems for modulation spaces. We completely characterize the continuity of a linear operator on the modulation spaces $M^p$ for every $1\leq p\leq\infty$, by the membership of its kernel to (mixed) modulation spaces.…
The purpose of this paper is to study the $L^p$ boundedness of operators of the form \[ f\mapsto \psi(x) \int f(\gamma_t(x))K(t)\: dt, \] where $\gamma_t(x)$ is a $C^\infty$ function defined on a neighborhood of the origin in $(t,x)\in…
Studying the stability of partially observed Markov decision processes (POMDPs) with respect to perturbations in either transition or observation kernels is a significant problem. While asymptotic robustness/stability results as approximate…
Much recent work in bioinformatics has focused on the inference of various types of biological networks, representing gene regulation, metabolic processes, protein-protein interactions, etc. A common setting involves inferring network edges…
We consider the problem of simultaneously learning to linearly combine a very large number of kernels and learn a good predictor based on the learnt kernel. When the number of kernels $d$ to be combined is very large, multiple kernel…