Related papers: An improved example for an autoconvolution inequal…
In this paper we consider a generalized version of Carleman's inequality. An equivalent version of it states that $\|f\|_{A_\alpha^{2\alpha}}\leq\|f\|_{H^2}$, where $f$ is a holomorphic function and $\alpha>1$. If the norms…
Recent work \cite{arifgroup} introduced Federated Proximal Gradient \textbf{(\texttt{FedProxGrad})} for solving non-convex composite optimization problems in group fair federated learning. However, the original analysis established…
Despite the empirical success of meta reinforcement learning (meta-RL), there are still a number poorly-understood discrepancies between theory and practice. Critically, biased gradient estimates are almost always implemented in practice,…
Can we accelerate convergence of gradient descent without changing the algorithm -- just by carefully choosing stepsizes? Surprisingly, we show that the answer is yes. Our proposed Silver Stepsize Schedule optimizes strongly convex…
Variational inequalities play a key role in machine learning research, such as generative adversarial networks, reinforcement learning, adversarial training, and generative models. This paper is devoted to the constrained variational…
We investigate approximation guarantees provided by logistic regression for the fundamental problem of agnostic learning of homogeneous halfspaces. Previously, for a certain broad class of "well-behaved" distributions on the examples,…
In this paper we compare two regression curves by measuring their difference by the area between the two curves, represented by their $L^1$-distance. We develop asymptotic confidence intervals for this measure and statistical tests to…
If the step distribution in a renewal process has finite mean and regularly varying tail with index -{\alpha}, 1<{\alpha}<2, the first two terms in the asymptotic expansion of the renewal function have been known for many years. Here we…
In this paper we consider the following problem of phase retrieval: Given a collection of real-valued band-limited functions $\{\psi_{\lambda}\}_{\lambda\in \Lambda}\subset L^2(\mathbb{R}^d)$ that constitutes a semi-discrete frame, we ask…
Phase retrieval is concerned with recovering a function $f$ from the absolute value of its Fourier transform $|\widehat{f}|$. We study the stability properties of this problem in Lebesgue spaces. Our main results shows that $$ \|…
In this paper, we study the convergence rate of the gradient (or steepest descent) method with fixed step lengths for finding a stationary point of an $L$-smooth function. We establish a new convergence rate, and show that the bound may be…
Adam is a popular variant of stochastic gradient descent for finding a local minimizer of a function. In the constant stepsize regime, assuming that the objective function is differentiable and non-convex, we establish the convergence in…
We analyze the properties of gradient descent on convex surrogates for the zero-one loss for the agnostic learning of linear halfspaces. If $\mathsf{OPT}$ is the best classification error achieved by a halfspace, by appealing to the notion…
We find two-sides estimates for the best uniform approximations of classes of convolutions of $2\pi$-periodic functions from unit ball of the space $L_p, 1 \le p <\infty,$ with fixed kernels, modules of Fourier coefficients of which satisfy…
We study the recovery of functions in various norms, including $L_p$ with $1\le p\le\infty$, based on function evaluations. We obtain worst case error bounds for general classes of functions in terms of the best $L_2$-approximation from a…
We study learning under a two-step contrastive example oracle, as introduced by Mansouri et. al. (2025), where each queried (or sampled) labeled example is paired with an additional contrastive example of opposite label. While Mansouri et…
An inequality for the $p$th power of the norm of a stochastic convolution integral in a Hilbert space is proved. The inequality is stronger than analogues inequalities in the Literature in the sense that it is pathwise and not in…
A significant milestone in modern gradient-based optimization was achieved with the development of Nesterov's accelerated gradient descent (NAG) method. This forward-backward technique has been further advanced with the introduction of its…
Motivated by the extensive application of approximate gradients in machine learning and optimization, we investigate inexact subgradient methods subject to persistent additive errors. Within a nonconvex semialgebraic framework, assuming…
Stochastic Gradient Descent (SGD) with adaptive steps is widely used to train deep neural networks and generative models. Most theoretical results assume that it is possible to obtain unbiased gradient estimators, which is not the case in…