Related papers: Sharp convex generalizations of stochastic Gronwal…
This work studies the generalization error of gradient methods. More specifically, we focus on how training steps $T$ and step-size $\eta$ might affect generalization in smooth stochastic convex optimization (SCO) problems. We first provide…
The stochastic subgradient method is a widely-used algorithm for solving large-scale optimization problems arising in machine learning. Often these problems are neither smooth nor convex. Recently, Davis et al. [1-2] characterized the…
Regularization is widely used in statistics and machine learning to prevent overfitting and gear solution towards prior information. In general, a regularized estimation problem minimizes the sum of a loss function and a penalty term. The…
We present a novel model Graph Neural Stochastic Differential Equations (Graph Neural SDEs). This technique enhances the Graph Neural Ordinary Differential Equations (Graph Neural ODEs) by embedding randomness into data representation using…
Recently there are a considerable amount of work devoted to the study of the algorithmic stability and generalization for stochastic gradient descent (SGD). However, the existing stability analysis requires to impose restrictive assumptions…
In deep latent Gaussian models, the latent variable is generated by a time-inhomogeneous Markov chain, where at each time step we pass the current state through a parametric nonlinear map, such as a feedforward neural net, and add a small…
We consider random perturbations of discrete-time dynamical systems. We give sufficient conditions for the stochastic stability of certain classes of maps, in a strong sense. This improves the main result in J. F. Alves, V. Araujo, Random…
We derive new diffusion solutions to the monoenergetic generalized linear Boltzmann transport equation (GLBE) for the stationary collision density and scalar flux about an isotropic point source in an infinite $d$-dimensional absorbing…
In this paper, we study an ordinary differential equation with a degenerate global attractor at the origin, to which we add a white noise with a small parameter that regulates its intensity. Under general conditions, for any fixed…
An influential line of recent work has focused on the generalization properties of unregularized gradient-based learning procedures applied to separable linear classification with exponentially-tailed loss functions. The ability of such…
Many large-scale and distributed optimization problems can be brought into a composite form in which the objective function is given by the sum of a smooth term and a nonsmooth regularizer. Such problems can be solved via a proximal…
For the discretization of the convective term in the Navier-Stokes equations (NSEs), the commonly used convective formulation (CONV) does not preserve the energy if the divergence constraint is only weakly enforced. In this paper, we apply…
We present a rough path analog of the classical Gronwall Lemma introduced recently by A. Deya, M. Gubinelli, M. Hofmanov\'a, S. Tindel in [arXiv:1604.00437] and discuss two of its applications. First, it is applied in the framework of rough…
We propose and analyze several stochastic gradient algorithms for finding stationary points or local minimum in nonconvex, possibly with nonsmooth regularizer, finite-sum and online optimization problems. First, we propose a simple proximal…
We consider stochastic gradient descent and its averaging variant for binary classification problems in a reproducing kernel Hilbert space. In the traditional analysis using a consistency property of loss functions, it is known that the…
Spatial differentiability of solutions of stochastic differential equations (SDEs) is a classical question in stochastic analysis. The case of coefficients with globally Lipschitz continuous derivatives is well understood in the literature.…
Due to the non-smoothness of optimization problems in Machine Learning, generalized smoothness assumptions have been gaining a lot of attention in recent years. One of the most popular assumptions of this type is $(L_0,L_1)$-smoothness…
A dynamic sampled stochastic approximated (DS-SA) extragradient method for stochastic variational inequalities (SVI) is proposed that is \emph{robust} with respect to an unknown Lipschitz constant $L$. To the best of our knowledge, it is…
We provide a new convergence analysis of stochastic gradient Langevin dynamics (SGLD) for sampling from a class of distributions that can be non-log-concave. At the core of our approach is a novel conductance analysis of SGLD using an…
Feature attributions are a popular tool for explaining the behavior of Deep Neural Networks (DNNs), but have recently been shown to be vulnerable to attacks that produce divergent explanations for nearby inputs. This lack of robustness is…