Related papers: A gradient estimate for the linearized translator …
We first show that any $4$-dimensional non-Ricci-flat steady gradient Ricci soliton singularity model must satisfy $|Rm|\leq cR$ for some positive constant $c$. Then, we apply the Hamilton-Ivey estimate to prove a quantitative lower bound…
Motivated by the testing condition for Radon-Brascamp-Lieb multilinear functionals established in arXiv:2201.12201, this paper is concerned with identifying local conditions on smooth maps $u(t)$ with values in the space of decomposable…
Velocity gradient tensor, $A_{ij}\equiv \partial u_i/\partial x_j$, in a turbulence flow field is modeled by separating the treatment of intermittent magnitude ($A = \sqrt{A_{ij}A_{ij}}$) from that of the more universal normalized velocity…
In distributed optimization and distributed numerical linear algebra, we often encounter an inversion bias: if we want to compute a quantity that depends on the inverse of a sum of distributed matrices, then the sum of the inverses does not…
The recent introduction of the gradient flow has provided a new tool to probe the dynamics of quantum field theories. The latest developments have shown how to use the gradient flow for the exploration of symmetries, and the definition of…
We find a near detailed balance solution to the relativistic Boltzmann equation under the relaxation time approximation with a collision term which differs from the Anderson-Witting model and is dependent on the stationary observer. Using…
In this paper we present new theoretical results for the Dantzig and Lasso estimators of the drift in a high dimensional Ornstein-Uhlenbeck model under sparsity constraints. Our focus is on oracle inequalities for both estimators and error…
The scalar curvature for the noncommutative four torus $\mathbb{T}_\Theta^4$, where its flat geometry is conformally perturbed by a Weyl factor, is computed by making the use of a noncommutative residue that involves integration over the…
We define new differential structures on the Wasserstein spaces $\mathcal{W}_p(M)$ for $p > 2$ and a general Riemannian manifold $(M,g)$. We consider a very general and possibly degenerate second order partial differential flow equation…
We quantify, uniformly over time and with high probability, the discrepancy between the predictions of a two-layer neural network trained by stochastic gradient descent (SGD) and their mean-field limit, for quadratic loss and ridge…
We present an estimate of the Wasserstein distance between the data distribution and the generation of score-based generative models. The sampling complexity with respect to dimension is $\mathcal{O}(\sqrt{d})$, with a logarithmic constant.…
Learning disentangled representations of textual data is essential for many natural language tasks such as fair classification, style transfer and sentence generation, among others. The existent dominant approaches in the context of text…
We study high-dimensional Bayesian linear regression with product priors. Using the nascent theory of non-linear large deviations (Chatterjee and Dembo,2016), we derive sufficient conditions for the leading-order correctness of the naive…
This paper derives non-asymptotic error bounds for nonlinear stochastic approximation algorithms in the Wasserstein-$p$ distance. To obtain explicit finite-sample guarantees for the last iterate, we develop a coupling argument that compares…
Motivated by recent developments in the fields of large deviations for interacting particle system and mean field control, we establish a comparison principle for the Hamilton--Jacobi equation corresponding to linearly controlled gradient…
In this paper, we provide the first precise distributional characterization of gradient descent iterates for general multi-layer neural networks under the canonical single-index regression model, in the `finite-width proportional regime'…
We introduce the Wasserstein Transform (WT), a general unsupervised framework for updating distance structures on given data sets with the purpose of enhancing features and denoising. Our framework represents each data point by a…
Wasserstein barycenters provide a principled approach for aggregating probability measures, while preserving the geometry of their ambient space. Existing discrete methods are not scalable as they assume access to the complete set of…
Let $(M,g)$ be a complete non-compact Riemannian manifold with the $m$-dimensional Bakry-\'{E}mery Ricci curvature bounded below by a non-positive constant. In this paper, we give a localized Hamilton-type gradient estimate for the positive…
The local Rademacher complexity framework is one of the most successful general-purpose toolboxes for establishing sharp excess risk bounds for statistical estimators based on the framework of empirical risk minimization. Applying this…