Related papers: Error Analysis of Generalized Langevin Equations w…
We study a class of semi-linear differential Volterra equations with polynomial-type potentials that incorporates the effects of memory while being subjected to random perturbations via an additive Gaussian noise. We show that for a broad…
We provide a framework to analyze the convergence of discretized kinetic Langevin dynamics for $M$-$\nabla$Lipschitz, $m$-convex potentials. Our approach gives convergence rates of $\mathcal{O}(m/M)$, with explicit stepsize restrictions,…
Subgradient algorithms for training support vector machines have been quite successful for solving large-scale and online learning problems. However, they have been restricted to linear kernels and strongly convex formulations. This paper…
The underdamped, non-linear, generalized Langevin equation is widely used to model coarse-grained dynamics of soft and biological materials. By means of a projection operator formalism, we show under which approximations this equation can…
Gaussian Processes (GPs) provide powerful probabilistic frameworks for interpolation, forecasting, and smoothing, but have been hampered by computational scaling issues. Here we investigate data sampled on one dimension (e.g., a scalar or…
This paper investigates the limit distribution of discretization errors in stochastic Volterra equations (SVEs) with general multidimensional kernel structures. While prior studies, such as Fukasawa and Ugai (2023), were focused on…
In this paper we investigate how gradient-based algorithms such as gradient descent, (multi-pass) stochastic gradient descent, its persistent variant, and the Langevin algorithm navigate non-convex loss-landscapes and which of them is able…
The irreversible generalized Langevin equation (iGLE) contains a nonstationary friction kernel that in certain limits reduces to the GLE with space-dependent friction. For more general forms of the friction kernel, the iGLE was previously…
Many algorithms in computer vision and robotics make strong assumptions about uncertainty, and rely on the validity of these assumptions to produce accurate and consistent state estimates. In practice, dynamic environments may degrade…
We propose a new theoretical framework that exploits convolution kernels to transform a Volterra-type path-dependent (non-Markovian) stochastic process into a standard (Markovian) diffusion process. Remarkably, it is also possible to go…
Perturbing a system far away from equilibrium via a time dependent protocol can formally be described by a nonlinear Volterra series expansion. Here we derive identities for the nonlinear memory kernels arising in such nonlinear expansion,…
While momentum-based accelerated variants of stochastic gradient descent (SGD) are widely used when training machine learning models, there is little theoretical understanding on the generalization error of such methods. In this work, we…
In this work, we unify several expected generalization error bounds based on random subsets using the framework developed by Hellstr\"om and Durisi [1]. First, we recover the bounds based on the individual sample mutual information from Bu…
The algorithms used to train neural networks, like stochastic gradient descent (SGD), have close parallels to natural processes that navigate a high-dimensional parameter space -- for example protein folding or evolution. Our study uses a…
This paper explores the residual based a posteriori error estimations for the generalized Burgers-Huxley equation (GBHE) featuring weakly singular kernels. Initially, we present a reliable and efficient error estimator for both the…
We investigate the generalization and optimization properties of shallow neural-network classifiers trained by gradient descent in the interpolating regime. Specifically, in a realizable scenario where model weights can achieve arbitrarily…
This paper is devoted to establishing the full scaling limit theorems for multivariate Hawkes processes. Under some mild conditions on the exciting kernels, we develop a new way to prove that after a suitable time-spatial scaling, the…
In Hezaveh et al. 2017 we showed that deep learning can be used for model parameter estimation and trained convolutional neural networks to determine the parameters of strong gravitational lensing systems. Here we demonstrate a method for…
The generalized Langevin equation is used as a model for various coarse-grained physical processes, e.g., the time evolution of the velocity of a given larger particle in an implicitly represented solvent, when the relevant time scales of…
Recently, several works have shown that natural modifications of the classical conditional gradient method (aka Frank-Wolfe algorithm) for constrained convex optimization, provably converge with a linear rate when: i) the feasible set is a…