Related papers: Neural Discovery of Strichartz Extremizers
We establish nearly optimal upper and lower bounds for approximating decision tree splits in data streams. For regression with labels in the range $\{0,1,\ldots,M\}$, we give a one-pass algorithm using $\tilde{O}(M^2/\epsilon)$ space that…
The need of fast distributed solvers for optimization problems in networked systems has motivated the recent development of the Fast-Lipschitz optimization framework. In such an optimization, problems satisfying certain qualifying…
This paper investigates maximizers of the information divergence from an exponential family $E$. It is shown that the $rI$-projection of a maximizer $P$ to $E$ is a convex combination of $P$ and a probability measure $P_-$ with disjoint…
Non-stationary approximations of the final value of a converging sequence are discussed, and we show that extremal eigenvalues can be reasonably estimated from the CG iterates without much computation at all. We introduce estimators of…
A precise characterization of the extremal points of sublevel sets of nonsmooth penalties provides both detailed information about minimizers, and optimality conditions in general classes of minimization problems involving them. Moreover,…
We address the problem of finding worst-case nonparametric bounds for T-statistic by considering the extremal problem of maximising the mid-quantile (a special case of 'smoothed quantile' as discussed in \cite{St77} and \cite{W11}) $\tilde…
This paper aims to give a general (possibly compact or noncompact) analog of Strichartz inequalities with loss of derivatives, obtained by Burq, G\'erard, and Tzvetkov [19] and Staffilani and Tataru [51]. Moreover we present a new approach,…
This paper introduces new parameterizations of equilibrium neural networks, i.e. networks defined by implicit equations. This model class includes standard multilayer and residual networks as special cases. The new parameterization admits a…
In this article, we establish the existence of an extremal function for the k-th order critical Hardy-Sobolev-Maz'ya (HSM) inequalities on the upper half space $\mathbb{R}^{n+1}_{+}$ when $k\ge 2$ and $n\geq 2k+2$:…
This paper studies the universal approximation property of deep neural networks for representing probability distributions. Given a target distribution $\pi$ and a source distribution $p_z$ both defined on $\mathbb{R}^d$, we prove under…
We present a simple and flexible method to prove consistency of semidefinite optimization problems on random graphs. The method is based on Grothendieck's inequality. Unlike the previous uses of this inequality that lead to constant…
In this paper, we prove convergence rates for time discretisation schemes for semi-linear stochastic evolution equations with additive or multiplicative Gaussian noise, where the leading operator $A$ is the generator of a strongly…
The Hausdorff-Young inequality for Euclidean space, in its sharp form due to Beckner, gives an upper bound for the Fourier transform in terms of Lebesgue space norms, with an optimal constant. The extremizers have been identified by Lieb to…
Solving partial differential equations (PDEs) is a central task in scientific computing. Recently, neural network approximation of PDEs has received increasing attention due to its flexible meshless discretization and its potential for…
As an alternative to PINNs, a Deep Ritz framework is proposed to solve fully nonlinear PDEs. A least-squares algorithm is advocated to decouple the nonlinearities from the variational features of several fully nonlinear PDEs. A splitting…
The convergence theory for the gradient sampling algorithm is extended to directionally Lipschitz functions. Although directionally Lipschitz functions are not necessarily locally Lipschitz, they are almost everywhere differentiable and…
We study asymmetric rank-one spiked tensor models in the high-dimensional regime, where the noise entries are independent and identically distributed with zero mean, unit variance, and finite fourth moment. This extends the classical…
Modern convolutional networks, incorporating rectifiers and max-pooling, are neither smooth nor convex; standard guarantees therefore do not apply. Nevertheless, methods from convex optimization such as gradient descent and Adam are widely…
Distributed optimization algorithms have been studied extensively in the literature; however, underlying most algorithms is a linear consensus scheme, i.e. averaging variables from neighbors via doubly stochastic matrices. We consider…
The recently developed average-case analysis of optimization methods allows a more fine-grained and representative convergence analysis than usual worst-case results. In exchange, this analysis requires a more precise hypothesis over the…