Related papers: An improved example for an autoconvolution inequal…
In the present paper, we formulate two versions of Frank--Wolfe algorithm or conditional gradient method to solve the DC optimization problem with an adaptive step size. The DC objective function consists of two components; the first is…
We prove the $\Gamma$-convergence of sequences of differentially constrained, random integral functionals of the form \begin{equation*} \int_{U} f\Big(\omega, x/\varepsilon, \mathbb{A} u\Big) \mathrm{d} x \end{equation*} for the class of…
A halfspace is a function $f\colon\{-1,1\}^n \rightarrow \{0,1\}$ of the form $f(x)=\mathbb{1}(a\cdot x>t)$, where $\sum_i a_i^2=1$. We show that if $f$ is a halfspace with $\mathbb{E}[f]=\epsilon$ and $a'=\max_i |a_i|$, then the degree-1…
We propose \textbf{ULU}, a novel non-monotonic, piecewise activation function defined as $\{f(x;\alpha_1),x<0; f(x;\alpha_2),x>=0 \}$, where $f(x;\alpha)=0.5x(tanh(\alpha x)+1),\alpha >0$. ULU treats positive and negative inputs…
We show that accelerated gradient descent, averaged gradient descent and the heavy-ball method for non-strongly-convex problems may be reformulated as constant parameter second-order difference equation algorithms, where stability of the…
We study the convolution function $$ C[f(x)] := \int_1^x f(y)f({x\over y}) {{\rm d} y\over y} $$ when $f(x)$ is a suitable number-theoretic error term. Asymptotics and upper bounds for $C[f(x)]$ are derived from mean square bounds for…
We show that if an essentially arbitrary sequence supported on an interval containing $x$ integers, is convolved with a tiny Siegel-Walfisz-type sequence supported on an interval containing $\exp((\log x)^{\varepsilon})$ integers then the…
When smoothing a function $f$ via convolution with some kernel, it is often desirable to adapt the amount of smoothing locally to the variation of $f$. For this purpose, the constant smoothing coefficient of regular convolutions needs to be…
An algorithm is given for determining an optimal $b$-step approximation of weighted data, where the error is measured with respect to the $L_\infty$ norm. For data presorted by the independent variable the algorithm takes $\Theta(n + \log n…
We establish a sub-convexity estimate for Rankin-Selberg $L$-functions in the combined level aspect, using the circle method. If $p$ and $q$ are distinct prime numbers, $f$ and $g$ are non-exceptional newforms (modular or Maass) for the…
Longitudinal binary or count functional data are common in neuroscience, but are often too large to analyze with existing functional regression methods. We propose one-step penalized generalized estimating equations that supports…
This work concerns the study of the subdifferential of the integral functional $$ E_f(x)=\int_{T} f(t,x)d\mu(t), $$ where $f$ is a (not necessarily convex) normal integrand, $({T},\mathcal{A},\mu)$ is a $\sigma$-finite measure space, while…
We present a strikingly simple proof that two rules are sufficient to automate gradient descent: 1) don't increase the stepsize too fast and 2) don't overstep the local curvature. No need for functional values, no line search, no…
Suppose $F: \mathbb{R}^{N} \rightarrow [0, +\infty)$ be a convex function of class $C^{2}(\mathbb{R}^{N} \backslash \{0\})$ which is even and positively homogeneous of degree 1. We denote $\gamma_1=\inf\limits_{u\in W^{1,…
Zhang refined the classical Sobolev inequality $\|f\|_{L^{Np/(N-p)}} \lesssim \| \nabla f \|_{L^p}$, where $1\leq p \lt N$, by replacing $\|\nabla f\|_{L^p}$ with a smaller quantity invariant by unimodular affine transformations. The…
We consider the convolution model where i.i.d. random variables $X_i$ having unknown density $f$ are observed with additive i.i.d. noise, independent of the $X$'s. We assume that the density $f$ belongs to either a Sobolev class or a class…
One-step generative modeling has emerged as a leading approach to amortize the inference cost of diffusion and flow-matching models. Among distillation-free methods, MeanFlow training is notoriously unstable, with non-decreasing loss and…
For a real-valued non-negative and log-concave function we introduce a notion of difference function; the difference function represents a functional analog on the difference body of a convex body. We prove a sharp inequality which bounds…
A common way to estimate an unknown convex regression function $f_0: \Omega \subset \mathbb{R}^d \rightarrow \mathbb{R}$ from a set of $n$ noisy observations is to fit a convex function that minimizes the sum of squared errors. However,…
Adaptive algorithms like AdaGrad and AMSGrad are successful in nonconvex optimization owing to their parameter-agnostic ability -- requiring no a priori knowledge about problem-specific parameters nor tuning of learning rates. However, when…