Related papers: Modified Equations for Stochastic Optimization
We study solutions to backward differential equations that are driven hybridly by a deterministic discontinuous rough path $W$ of finite $q$-variation for $q \in [1, 2)$ and by Brownian motion $B$. To distinguish between integration of…
We present two stochastic descent algorithms that apply to unconstrained optimization and are particularly efficient when the objective function is slow to evaluate and gradients are not easily obtained, as in some PDE-constrained…
We present a numerical method for the approximation of solutions for the class of stochastic differential equations driven by Brownian motions which induce stochastic variation in fixed directions. This class of equations arises naturally…
The G-Brownian-motion-driven stochastic differential equations (G-SDEs) as well as the G-expectation, which were seminally proposed by Peng and his colleagues, have been extensively applied to describing a particular kind of uncertainty…
We introduce stochastic normalizing flows, an extension of continuous normalizing flows for maximum likelihood estimation and variational inference (VI) using stochastic differential equations (SDEs). Using the theory of rough paths, the…
Diffusion approximation provides weak approximation for stochastic gradient descent algorithms in a finite time horizon. In this paper, we introduce new tools motivated by the backward error analysis of numerical stochastic differential…
We consider a class of stochastic smooth convex optimization problems under rather general assumptions on the noise in the stochastic gradient observation. As opposed to the classical problem setting in which the variance of noise is…
We have identified a potential method for unifying first-order optimizers through the use of variable Second-Moment Exponential Scaling(SMES). We begin with back propagation, addressing classic phenomena such as gradient vanishing and…
This paper proposes a probabilistic motion prediction method for long motions. The motion is predicted so that it accomplishes a task from the initial state observed in the given image. While our method evaluates the task achievability by…
The Malliavin differentiability of a SDE plays a crucial role in the study of density smoothness and ergodicity among others. For Gaussian driven SDEs the differentiability property is now well established. In this paper, we consider the…
In this paper, we introduce a new simple approach to developing and establishing the convergence of splitting methods for a large class of stochastic differential equations (SDEs), including additive, diagonal and scalar noise types. The…
The scaled Brownian motion (SBM) is regarded as one of the paradigmatic random processes, featuring the anomalous diffusion property characterized by the diffusion exponent. It is a Gaussian, self-similar process with independent…
In recent years, an intensive study of strong approximation of stochastic differential equations (SDEs) with a drift coefficient that may have discontinuities in space has begun. In many of these results it is assumed that the drift…
In this paper we study different algorithms for backward stochastic differential equations (BSDE in short) basing on random walk framework for 1-dimensional Brownian motion. Implicit and explicit schemes for both BSDE and reflected BSDE are…
We derive high-dimensional scaling limits and fluctuations for the online least-squares Stochastic Gradient Descent (SGD) algorithm by taking the properties of the data generating model explicitly into consideration. Our approach treats the…
We consider the distributed optimization problem where $n$ agents each possessing a local cost function, collaboratively minimize the average of the $n$ cost functions over a connected network. Assuming stochastic gradient information is…
A framework is introduced for solving a sequence of slowly changing optimization problems, including those arising in regression and classification applications, using optimization algorithms such as stochastic gradient descent (SGD). The…
The non-asymptotic analysis of Stochastic Gradient Descent (SGD) typically yields bounds that decompose into a bias term and a variance term. In this work, we focus on the bias component and study the extent to which SGD can match the…
We unify and extend the semigroup and the PDE approaches to stochastic maximal regularity of time-dependent semilinear parabolic problems with noise given by a cylindrical Brownian motion. We treat random coefficients that are only…
Stochastic gradient descent (SGD) is a popular algorithm for minimizing objective functions that arise in machine learning. For constant step-sized SGD, the iterates form a Markov chain on a general state space. Focusing on a class of…