Related papers: An improved example for an autoconvolution inequal…
In neural network training, RMSProp and Adam remain widely favoured optimisation algorithms. One of the keys to their performance lies in selecting the correct step size, which can significantly influence their effectiveness. Additionally,…
In this paper, the optimal convergence rate $O\left(N^{-1/2}\right)$ (where $N$ is the total number of iterations performed by the algorithm), without the presence of a logarithmic factor, is proved for mirror descent algorithms with…
We study the asymptotic convergence of solutions as $t\rightarrow\infty$ of $\partial_t u=-f(u)+\int f(u)$, a nonlocal differential equation that is formally a gradient flow in a constant-mass subspace of $L^2$ arising from simplified…
This paper investigates some aspects of the variational behaviour of nonsmooth functions, with special emphasis on certain stability phenomena. Relationships linking such properties as sharp minimality, superstability, error bound and…
We consider optimizing a function smooth convex function $f$ that is the average of a set of differentiable functions $f_i$, under the assumption considered by Solodov [1998] and Tseng [1998] that the norm of each gradient $f_i'$ is bounded…
Constant step-size Stochastic Gradient Descent exhibits two phases: a transient phase during which iterates make fast progress towards the optimum, followed by a stationary phase during which iterates oscillate around the optimal point. In…
Let $f\in L_{2\pi}$ be a real-valued even function with its Fourier series $ \frac{a_{0}}{2}+\sum_{n=1}^{\infty}a_{n}\cos nx,$ and let $S_{n}(f,x), n\geq 1,$ be the $n$-th partial sum of the Fourier series. It is well-known that if the…
The Riesz-Sobolev inequality relates the convolution of nonnegative functions on Euclidean space to the convolution of their symmetric nonincreasing rearrangements. We show that for dimension one, for indicator functions of sets, if the…
We provide non-asymptotic bounds for the well-known temporal difference learning algorithm TD(0) with linear function approximators. These include high-probability bounds as well as bounds in expectation. Our analysis suggests that a…
We present here a new method for approximating functions defined on superreflexive Banach spaces by differentiable functions with $\alpha$-H\"older derivatives (for some $0<\alpha\leq 1$). The smooth approximation is given by means of an…
We show how to obtain improved active learning methods in the agnostic (adversarial noise) setting by combining marginal leverage score sampling with non-independent sampling strategies that promote spatial coverage. In particular, we…
Convolutional autoencoders have emerged as popular methods for unsupervised defect segmentation on image data. Most commonly, this task is performed by thresholding a pixel-wise reconstruction error based on an $\ell^p$ distance. This…
Let $\Omega \subset \mathbb{R}^n$ be a convex domain and let $f:\Omega \rightarrow \mathbb{R}$ be a positive, subharmonic function (i.e. $\Delta f \geq 0$). Then $$ \frac{1}{|\Omega|} \int_{\Omega}{f dx} \leq \frac{c_n}{ |\partial \Omega| }…
Structured constraints in Machine Learning have recently brought the Frank-Wolfe (FW) family of algorithms back in the spotlight. While the classical FW algorithm has poor local convergence properties, the Away-steps and Pairwise FW…
Current state-of-art feature-engineered and end-to-end Automated Essay Score (AES) methods are proven to be unable to detect adversarial samples, e.g. the essays composed of permuted sentences and the prompt-irrelevant essays. Focusing on…
The Bregman distance is a central tool in convex optimization, particularly in first-order gradient descent and proximal-based algorithms. Such methods enable optimization of functions without Lipschitz continuous gradients by leveraging…
We consider the model of nonregular nonparametric regression where smoothness constraints are imposed on the regression function $f$ and the regression errors are assumed to decay with some sharpness level at their endpoints. The aim of…
We analyze the constant step size subgradient method on nonsmooth, nonconvex functions. We identify geometric assumptions on the objective function under which i) its domain admits a partition (stratification) into smooth manifolds (strata)…
This paper develops tests for inequality constraints of nonparametric regression functions. The test statistics involve a one-sided version of $L_p$-type functionals of kernel estimators $(1 \leq p < \infty)$. Drawing on the approach of…
In this paper, we propose a simple, fast and easy to implement algorithm LOSSGRAD (locally optimal step-size in gradient descent), which automatically modifies the step-size in gradient descent during neural networks training. Given a…