Related papers: Optimal upper bounds on expected kth record values…
This paper gives upper and lower bounds on the gap in Jensen's inequality, i.e., the difference between the expected value of a function of a random variable and the value of the function at the expected value of the random variable. The…
We study the generalization properties of the popular stochastic optimization method known as stochastic gradient descent (SGD) for optimizing general non-convex loss functions. Our main contribution is providing upper bounds on the…
Let recall that the term 'k-th extreme' was introduced in a limiting sense. That is, if $X_{r:n}$ denote the r-th order statistic then for fix k, as $n\to\infty$, $X_{n-k+1:n}$ is called the k-th extremes or k-th largest order statistics.…
This work provides a framework to compute an upper bound on the robust peak-to-peak gain of discrete-time uncertain linear systems using integral quadratic constraints (IQCs). Such bounds are of particular interest in the computation of…
For a general class of non-negative functions defined on integral ideals of number fields, upper bounds are established for their average over the values of certain principal ideals that are associated to irreducible binary forms with…
We study the rate of growth of ergodic sums along a sequence (a_n) of times: S_N f(x)=f(T^{a_1}x) + ... + f(T^{a_N}x). We characterize the maximal rate of growth of these ergodic sums and identify a number of sequences such as (2^n) that…
The maximum entropy principle is a powerful tool for solving underdetermined inverse problems. This paper considers the problem of discretizing a continuous distribution, which arises in various applied fields. We obtain the approximating…
We study the problem of estimating a distribution over a finite alphabet from an i.i.d. sample, with accuracy measured in relative entropy (Kullback-Leibler divergence). While optimal bounds on the expected risk are known, high-probability…
Given a finite set of points $S\subset\mathbb{R}^d$, a $k$-set of $S$ is a subset $A \subset S$ of size $k$ which can be strictly separated from $S \setminus A $ by a hyperplane. Similarly, a $k$-facet of a point set $S$ in general position…
In clinical trials and other applications, we often see regions of the feature space that appear to exhibit interesting behaviour, but it is unclear whether these observed phenomena are reflected at the population level. Focusing on a…
In this paper, with the notion of independent identically distributed (IID) random variables under sublinear expectations introduced by Peng [7-9], we investigate moment bounds for IID sequences under sublinear expectations. We can obtain a…
We derive a general upper bound for the number of incidences with $k$-dimensional varieties in ${\mathbb R}^d$. The leading term of this new bound generalizes previous bounds for the special cases of $k=1, k=d-1,$ and $k= d/2$, to every…
We study problem-dependent rates, i.e., generalization errors that scale near-optimally with the variance, the effective loss, or the gradient norms evaluated at the "best hypothesis." We introduce a principled framework dubbed "uniform…
In this paper, we use tools from rate-distortion theory to establish new upper bounds on the generalization error of statistical distributed learning algorithms. Specifically, there are $K$ clients whose individually chosen models are…
Efficient global optimization is the problem of minimizing an unknown function f, using as few evaluations f(x) as possible. It can be considered as a continuum-armed bandit problem, with noiseless data and simple regret. Expected…
We study the performance of stochastic gradient descent (SGD) on smooth and strongly-convex finite-sum optimization problems. In contrast to the majority of existing theoretical works, which assume that individual functions are sampled with…
It is well known that under general regularity conditions the distribution of the maximum likelihood estimator (MLE) is asymptotically normal. Very recently, bounds of the optimal order $O(1/\sqrt n)$ on the closeness of the distribution of…
Let $X$ be a random variable and define its concentration function by $$\mathcal{Q}_{h}(X)=\sup_{x\in \mathbb{R}}\mathbb{P}(X\in (x,x+h]).$$ For a sum $S_n=X_1+\cdots+X_n$ of independent real-valued random variables the Kolmogorov-Rogozin…
Let $n\geq 1$, $K>0$, and let $X=(X_1,X_2,\dots,X_n)$ be a random vector in $\mathbb{R}^n$ with independent $K$--subgaussian components. We show that for every $1$--Lipschitz convex function $f$ in $\mathbb{R}^n$ (the Lipschitzness with…
KSG mutual information estimator, which is based on the distances of each sample to its k-th nearest neighbor, is widely used to estimate mutual information between two continuous random variables. Existing work has analyzed the convergence…