Related papers: Improving discrepancy by moving a few points
Despite of many measures applied for determine the difference between two groups of observations, such as mean value, median value, sample stan- dard deviation and so on, we propose a novel non parametric transformation method based on…
A method for moving least squares interpolation and differentiation is presented in the framework of orthogonal polynomials on discrete points. This yields a robust and efficient method which can avoid singularities and breakdowns in the…
We establish inequalities that compare the p-Wasserstein distance to distances which are built as suprema of box measures. More precisely, when the measures are supported on $[0,1]^d$, we obtain sharp upper-bounds of the $p$-Wasserstein…
Change point estimation in its offline version is traditionally performed by optimizing over the data set of interest, by considering each data point as the true location parameter and computing a data fit criterion. Subsequently, the data…
Low discrepancy point sets have been widely used as a tool to approximate continuous objects by discrete ones in numerical processes, for example in numerical integration. Following a century of research on the topic, it is still unclear…
We address the classical issue of appropriate choice of the regularization and discretization level for the Tikhonov regularization of an inverse problem with imperfectly measured data. We focus on the fact that the proper choice of the…
The most common way to sample from a probability distribution is to use Monte-Carlo methods. For distributions on a continuous state space, one can find diffusions with the target distribution as equilibrium measure, so that the state of…
Performance estimation under covariate shift is a crucial component of safe AI model deployment, especially for sensitive use-cases. Recently, several solutions were proposed to tackle this problem, most leveraging model predictions or…
Let $S_n$ be the set of permutations on $\{1,\,\dots,\,n\}$ and $\pi\in S_n$. Let $\mathrm{d}(\pi)$ be the arithmetic average of $\{|i-\pi(i)|;\;1\le i\le n\}$. Then $\mathrm{d}(\pi)/n\in[0,\,1/2]$, the expected value of $\mathrm{d}(\pi)/n$…
Out-of-Distribution (OOD) detection is essential for the trustworthiness of AI systems. Methods using prior information (i.e., subspace-based methods) have shown effective performance by extracting information geometry to detect OOD data…
In high dimension, low sample size (HDLSS)settings, the simple average distance classifier based on the Euclidean distance performs poorly if differences between the locations get masked by the scale differences. To rectify this issue,…
Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions. Applying these methods in practice can require…
For a graph $G$ spanning a metric space, the dilation of a pair of points is the ratio of their distance in the shortest path graph metric to their distance in the metric space. Given a graph $G$ and a budget $k$, a classic problem is to…
Consider linear regression where the examples are generated by an unknown distribution on $R^d\times R$. Without any assumptions on the noise, the linear least squares solution for any i.i.d. sample will typically be biased w.r.t. the least…
We present an efficient algorithm that, given a discrete random variable $X$ and a number $m$, computes a random variable whose support is of size at most $m$ and whose Kolmogorov distance from $X$ is minimal, also for the one-sided…
For all $s \geq 1$ and $N \geq 1$ there exist sequences $(z_1,\ldots,z_N)$ in $[0,1]^s$ such that the star-discrepancy of these points can be bounded by $$D_N^*(z_1,\ldots,z_N) \leq c \frac{\sqrt{s}}{\sqrt{N}}.$$ The best known value for…
Information distance is a parameter-free similarity measure based on compression, used in pattern recognition, data mining, phylogeny, clustering, and classification. The notion of information distance is extended from pairs to multiples…
In this paper, we propose a method for estimating the distribution of time differences between connected events (such as ad impressions and corresponding customer calls). A special feature of this method is that it does not require matching…
Comparing probability distributions is a fundamental problem in data sciences. Simple norms and divergences such as the total variation and the relative entropy only compare densities in a point-wise manner and fail to capture the geometric…
You measure the value of a quantity x for a number of systems (cells, molecules, people, chunks of metal, DNA vectors, etc.). You repeat the whole set of measures in different occasions or assays, which you try to design as equal to one…