Related papers: Learning Sums of Independent Random Variables with…
For a sample of absolutely bounded i.i.d. random variables with a continuous density the cumulative distribution function of the sample variance is represented by a univariate integral over a Fourier series. If the density is a polynomial…
We introduce the notion of a reproducible algorithm in the context of learning. A reproducible learning algorithm is resilient to variations in its samples -- with high probability, it returns the exact same output when run on two samples…
In this paper, we establish a sufficient condition to compare linear combinations of independent and identically distributed (iid) infinite-mean random variables under usual stochastic order. We introduce a new class of distributions that…
The aim of this paper is to estimate the density f of a random variable X when one has access to independent observations of the sum of K $\ge$ 2 independent copies of X. We provide a constructive estimator based on a suitable definition of…
We study the full distribution $P_{N}\left(A\right)$ of sums $A = \sum_{i=1}^N$ where $x_1, \dots, x_N$ are $N \gg 1$ independent and identically distributed random variables each sampled from a given distribution $p(x)$ with a…
In the setting of entangled single-sample distributions, the goal is to estimate some common parameter shared by a family of distributions, given one \emph{single} sample from each distribution. We study mean estimation and linear…
Many machine learning algorithms are based on the assumption that training examples are drawn independently. However, this assumption does not hold anymore when learning from a networked sample because two or more training examples may…
We propose a new approach for metric learning by framing it as learning a sparse combination of locally discriminative metrics that are inexpensive to generate from the training data. This flexible framework allows us to naturally derive…
We consider nonlinear, or "event-dependent", sampling, i.e. such that the sampling instances {tk} depend on the function being sampled. The use of such sampling in the construction of Lebesgue's integral sums is noted and discussed as…
This paper proposes a comprehensive and unprecedented framework that streamlines the derivation of exact, compact -- yet tractable -- solutions for the probability density function (PDF) and cumulative distribution function (CDF) of the sum…
Let $\mathcal{A}$ be a sequence of $rk$ terms which is made up of $k$ distinct integers each appearing exactly $r$ times in $\mathcal{A}$. The sum of all terms of a subsequence of $\mathcal{A}$ is called a subsequence sum of $\mathcal{A}$.…
As predictive algorithms grow in popularity, using the same dataset to both train and test a new model has become routine across research, policy, and industry. Sample-splitting attains valid inference on model properties by using separate…
The distribution of sum of independent non-identical binomial random variables is frequently encountered in areas such as genomics, healthcare, and operations research. Analytical solutions to the density and distribution are usually…
We derive an information-theoretic lower bound for sample complexity in sparse recovery problems where inputs can be chosen sequentially and adaptively. This lower bound is in terms of a simple mutual information expression and unifies many…
AI systems typically make decisions and find patterns in data based on the computation of aggregate and specifically sum functions, expressed as queries, on data's attributes. This computation can become costly or even inefficient when…
We give a comparison inequality that allows one to estimate the tail probabilities of sums of independent Banach space valued random variables in terms of those of independent identically distributed random variables. More precisely, let…
We revisit the problem of distribution learning within the framework of learning-augmented algorithms. In this setting, we explore the scenario where a probability distribution is provided as potentially inaccurate advice on the true,…
Let X be a second countable locally compact Abelian group. Let $\xi_1, \xi_2$ be independent random variables with values in the group X and distributions $\mu_1, \mu_2$ such that the sum $\xi_1+\xi_2$ and the difference $\xi_1-\xi_2$ are…
Adaptivity is an important feature of data analysis---the choice of questions to ask about a dataset often depends on previous interactions with the same dataset. However, statistical validity is typically studied in a nonadaptive model,…
We derive a well-defined renormalized version of mutual information that allows to estimate the dependence between continuous random variables in the important case when one is deterministically dependent on the other. This is the situation…