Related papers: An improved example for an autoconvolution inequal…
We present an algorithm for testing halfspaces over arbitrary, unknown rotation-invariant distributions. Using $\tilde O(\sqrt{n}\epsilon^{-7})$ random examples of an unknown function $f$, the algorithm determines with high probability…
Motivation: Alignment-free (AF) distance/similarity functions are a key tool for sequence analysis. Experimental studies on real datasets abound and, to some extent, there are also studies regarding their control of false positive rate…
We develop and analyze a variant of Nesterov's accelerated gradient descent (AGD) for minimization of smooth non-convex functions. We prove that one of two cases occurs: either our AGD variant converges quickly, as if the function was…
We present the Trust Region Adversarial Functional Subdifferential (TRAFS) algorithm for constrained optimization of nonsmooth convex Lipschitz functions. Unlike previous methods that assume a subgradient oracle model, we work with the…
Let $\mu$ be a non-atomic self-similar measure on $\mathbb{R}$, and let $\nu$ be its pushforward to a non-degenerate curve in $\mathbb{R}^d, d\geq 1$. We show that for every $\epsilon>0$, there is $p>1$, so that $\left \lVert \hat{\nu}…
Let $M$ denote the centered Hardy--Littlewood operator on $\mathbb{R}$. We prove that \[ {\rm Var} (Mf)\le {\rm Var} (f) - \frac12\big| |f(\infty)|-|f(-\infty)|\big| \] for piecewise constant functions $f$ with nonzero and zero values…
In the context of global optimization of mixed-integer nonlinear optimization formulations, we consider smoothing univariate functions $f$ that satisfy $f(0)=0$, $f$ is increasing and concave on $[0,+\infty)$, $f$ is twice differentiable on…
We study when the \emph{optimization curve} of first-order methods -- the sequence \${f(x\_n)}*{n\ge0}\$ produced by constant-stepsize iterations -- is convex, equivalently when the forward differences \$f(x\_n)-f(x*{n+1})\$ are…
To address the difficult problem of multi-step ahead prediction of non-parametric autoregressions, we consider a forward bootstrap approach. Employing a local constant estimator, we can analyze a general type of non-parametric time series…
We show that standard deviation $\s$ satisfies the Leibniz inequality $\s(fg) \leq \s(f)\|g\| + \|f\|\s(g)$ for bounded functions f, g on a probability space, where the norm is the supremum norm. A related inequality that we refer to as…
We show that for any $k$-times continuously differentiable function $f:[a,\infty)\longrightarrow{\mathbb R}$, any integer $q\ge 0$ and any $\alpha>1$ the inequality $$\liminf_{x\to\infty} \frac{x^k \cdot\log x\cdot \log_2 x\cdot\dots\cdot…
Stochastic gradient descent is the method of choice for large scale optimization of machine learning objective functions. Yet, its performance is greatly variable and heavily depends on the choice of the stepsizes. This has motivated a…
Given a non-convex twice differentiable cost function f, we prove that the set of initial conditions so that gradient descent converges to saddle points where \nabla^2 f has at least one strictly negative eigenvalue has (Lebesgue) measure…
We study the convergence properties of the original and away-step Frank-Wolfe algorithms for linearly constrained stochastic optimization assuming the availability of unbiased objective function gradient estimates. The objective function is…
We show that alpha stable L\'evy motions can be simulated by any ergodic and aperiodic probability preserving transformation. Namely we show: - for $0<\alpha<1$ and every $\alpha$ stable L\'evy motion $\mathbb{W}$, there exists a function f…
We prove convergence rates of Stochastic Zeroth-order Gradient Descent (SZGD) algorithms for Lojasiewicz functions. The SZGD algorithm iterates as \begin{align*} \mathbf{x}_{t+1} = \mathbf{x}_t - \eta_t \widehat{\nabla} f (\mathbf{x}_t),…
We consider $L^2$-approximation on weighted reproducing kernel Hilbert spaces of functions depending on infinitely many variables. We focus on unrestricted linear information, admitting evaluations of arbitrary continuous linear…
We discuss three convolution inequalities that are connected to additive combinatorics. Cloninger and the second author showed that for nonnegative $f \in L^1(-1/4, 1/4)$, $$ \max_{-1/2 \leq t \leq 1/2} \int_{\mathbb{R}}{f(t-x) f(x) dx}…
Recent papers have shown that the Frank-Wolfe algorithm (FW) with open-loop step-sizes exhibits rates of convergence faster than the iconic $\mathcal{O}(t^{-1})$ rate. In particular, when the minimizer of a strongly convex function over a…
This paper proposes a new estimation procedure for the ambiguity function of a non-stationary time series. The stochastic properties of the empirical ambiguity function calculated from a single sample in time are derived. Different…