Related papers: On closeness to k-wise uniformity
We consider the problem of hypothesis testing for discrete distributions. In the standard model, where we have sample access to an underlying distribution $p$, extensive research has established optimal bounds for uniformity testing,…
Quantization for a probability distribution refers to the idea of estimating a given probability by a discrete probability supported by a finite number of points. In this paper, firstly a general approach to this process is outlined using…
We consider one-dimensional discrete-time random walks (RWs) of $n$ steps, starting from $x_0=0$, with arbitrary symmetric and continuous jump distributions $f(\eta)$, including the important case of L\'evy flights. We study the statistics…
We show that for every $k\in\mathbb{N}$ and $\varepsilon>0$, for large enough alphabet $R$, given a $k$-CSP with alphabet size $R$, it is NP-hard to distinguish between the case that there is an assignment satisfying at least…
The following generalisation of the Erd\H{o}s unit distance problem was recently suggested by Palsson, Senger and Sheffer. Given $k$ positive real numbers $\delta_1,\dots,\delta_k$, a $(k+1)$-tuple $(p_1,\dots,p_{k+1})$ in $\mathbb{R}^d$ is…
In this paper, we obtain quantitative, non-asymptotic, and data-dependent \textit{Bernstein-von Mises type} bounds on the normal approximation of the posterior distribution in exponential family models with arbitrary centring and scaling.…
The $k$-nearest neighbour ($k$-NN) classifier is one of the oldest and most important supervised learning algorithms for classifying datasets. Traditionally the Euclidean norm is used as the distance for the $k$-NN classifier. In this…
This work presents an upper-bound to value that the Kullback-Leibler (KL) divergence can reach for a class of probability distributions called quantum distributions (QD). The aim is to find a distribution $U$ which maximizes the KL…
We consider the problem of estimating the edge density of densest $K$-node subgraphs of an Erd\"os-R\'{e}nyi graph $\mathbb{G}(n,1/2)$. The problem is well-understood in the regime $K=\Theta(\log n)$ and in the regime $K=\Theta(n)$. In the…
We investigate the statistical task of closeness (or equivalence) testing for multidimensional distributions. Specifically, given sample access to two unknown distributions $\mathbf p, \mathbf q$ on $\mathbb R^d$, we want to distinguish…
We study the problem of testing whether an unknown $n$-variable Boolean function is a $k$-junta in the distribution-free property testing model, where the distance between functions is measured with respect to an arbitrary and unknown…
In this note we establish a uniform bound for the distribution of a sum $S_n=X_1+\cdots+X_n$ of independent non-homogeneous Bernoulli trials. Specifically, we prove that $\sigma_n \mathbb{P}(S_n\!=\!j)\leq\eta$ where $\sigma_n$ denotes the…
For probability distributions on $\mathbb{R}^n$, we study the optimal sample size N = N(n,p) that suffices to uniformly approximate the pth moments of all one-dimensional marginals. Under the assumption that the marginals have bounded 4p…
Distribution testing can be described as follows: $q$ samples are being drawn from some unknown distribution $P$ over a known domain $[n]$. After the sampling process, a decision must be made about whether $P$ holds some property, or is far…
We study the maximum $k$-set coverage problem in the following distributed setting. A collection of sets $S_1,\ldots,S_m$ over a universe $[n]$ is partitioned across $p$ machines and the goal is to find $k$ sets whose union covers the most…
A graph $G$ is said to be $\mathcal H(n,\Delta)$-universal if it contains every graph on $n$ vertices with maximum degree at most $\Delta$. It is known that for any $\varepsilon > 0$ and any natural number $\Delta$ there exists $c > 0$ such…
A Poisson Binomial distribution over $n$ variables is the distribution of the sum of $n$ independent Bernoullis. We provide a sample near-optimal algorithm for testing whether a distribution $P$ supported on $\{0,...,n\}$ to which we have…
Given a discrete-valued sample $X_1,...,X_n$ we wish to decide whether it was generated by a distribution belonging to a family $H_0$, or it was generated by a distribution belonging to a family $H_1$. In this work we assume that all…
We give a highly efficient "semi-agnostic" algorithm for learning univariate probability distributions that are well approximated by piecewise polynomial density functions. Let $p$ be an arbitrary distribution over an interval $I$ which is…
Let $0<m<n$ be integers, and let $K_w$ denote the completion of a number field $K$ at a non-trivial place $w$. For each non-zero $\textbf{u}\in K_w^n$, let $\omega_{m-1}(\textbf{u})$ denote the exponent of best approximation to $\textbf{u}$…