English
Related papers

Related papers: Deep BSDE Solver on Bounded Domains Part I: Genera…

200 papers

The $O(1/k^2)$ convergence rate in function value of accelerated gradient descent is optimal, but there are many modifications that have been used to speed up convergence in practice. Among these modifications are restarts, that is,…

Optimization and Control · Mathematics 2023-10-12 Walaa M. Moursi , Viktor Pavlovic , Stephen A. Vavasis

Deep learning has aroused extensive attention due to its great empirical success. The efficiency of the block coordinate descent (BCD) methods has been recently demonstrated in deep neural network (DNN) training. However, theoretical…

Optimization and Control · Mathematics 2019-05-14 Jinshan Zeng , Tim Tsz-Kit Lau , Shaobo Lin , Yuan Yao

The aim of this work is to propose an extension of the deep solver by Han, Jentzen, E (2018) to the case of forward backward stochastic differential equations (FBSDEs) with jumps. As in the aforementioned solver, starting from a discretized…

Probability · Mathematics 2025-05-23 Kristoffer Andersson , Alessandro Gnoatto , Marco Patacca , Athena Picarelli

Generalization bounds which assess the difference between the true risk and the empirical risk have been studied extensively. However, to obtain bounds, current techniques use strict assumptions such as a uniformly bounded or a Lipschitz…

Machine Learning · Computer Science 2020-02-25 Yossi Adi , Yaniv Nemcovsky , Alex Schwing , Tamir Hazan

We study the overparametrization bounds required for the global convergence of stochastic gradient descent algorithm for a class of one hidden layer feed-forward neural networks, considering most of the activation functions used in…

Machine Learning · Computer Science 2022-11-17 Bartłomiej Polaczyk , Jacek Cyranka

The paper establishes the strong convergence rates of a spatio-temporal full discretization of the stochastic wave equation with nonlinear damping in dimension one and two. We discretize the SPDE by applying a spectral Galerkin method in…

Numerical Analysis · Mathematics 2024-12-30 Meng Cai , David Cohen , Xiaojie Wang

Classifiers built with neural networks handle large-scale high dimensional data, such as facial images from computer vision, extremely well while traditional statistical methods often fail miserably. In this paper, we attempt to understand…

Machine Learning · Statistics 2020-02-04 Tianyang Hu , Zuofeng Shang , Guang Cheng

We consider statistical tasks in high dimensions whose loss depends on the data only through its projection into a fixed-dimensional subspace spanned by the parameter vectors and certain ground truth vectors. This includes classifying…

Machine Learning · Statistics 2025-12-23 Reza Gheissari , Aukosh Jagannath

The first order loss function and its complementary function are extensively used in practical settings. When the random variable of interest is normally distributed, the first order loss function can be easily expressed in terms of the…

Optimization and Control · Mathematics 2014-09-09 Roberto Rossi , S. Armagan Tarim , Steven Prestwich , Brahim Hnich

Stochastic gradient descent (SGD) has been a go-to algorithm for nonconvex stochastic optimization problems arising in machine learning. Its theory however often requires a strong framework to guarantee convergence properties. We hereby…

Optimization and Control · Mathematics 2025-03-11 Azar Louzi

We study the Finite-Dimensional Distributions (FDDs) of deep neural networks with randomly initialized weights that have finite-order moments. Specifically, we establish Gaussian approximation bounds in the Wasserstein-$1$ norm between the…

Machine Learning · Statistics 2026-03-05 Krishnakumar Balasubramanian , Nathan Ross

We propose an efficient distributed randomized coordinate descent method for minimizing regularized non-strongly convex loss functions. The method attains the optimal $O(1/k^2)$ convergence rate, where $k$ is the iteration counter. The core…

Optimization and Control · Mathematics 2014-07-29 Olivier Fercoq , Zheng Qu , Peter Richtárik , Martin Takáč

In this paper, we study a kind of constrained backward stochastic differential equations (BSDEs) such that the nonlinear expectation of the composition of a loss function and the solution remains above zero. The existence and uniqueness…

Probability · Mathematics 2025-11-24 Hanwu Li

Density ratio estimation (DRE) is a core technique in machine learning used to capture relationships between two probability distributions. $f$-divergence loss functions, which are derived from variational representations of $f$-divergence,…

Machine Learning · Computer Science 2025-03-18 Yoshiaki Kitazawa

Stochastic partial differential equations (SPDEs) are ubiquitous in engineering and computational sciences. The stochasticity arises as a consequence of uncertainty in input parameters, constitutive relations, initial/boundary conditions,…

Data Analysis, Statistics and Probability · Physics 2020-01-29 Sharmila Karumuri , Rohit Tripathy , Ilias Bilionis , Jitesh Panchal

The Sinc approximation is a function approximation formula that attains exponential convergence for rapidly decaying functions defined on the whole real axis. Even for other functions, the Sinc approximation works accurately when combined…

Numerical Analysis · Computer Science 2022-03-04 Tomoaki Okayama

We prove linear convergence of gradient descent to a global optimum for the training of deep residual networks with constant layer width and smooth activation function. We show that if the trained weights, as a function of the layer index,…

Machine Learning · Computer Science 2023-01-26 Rama Cont , Alain Rossier , RenYuan Xu

Loss functions are at the heart of deep learning, shaping how models learn and perform across diverse tasks. They are used to quantify the difference between predicted outputs and ground truth labels, guiding the optimization process to…

Machine Learning · Computer Science 2025-09-11 Omar Elharrouss , Yasir Mahmood , Yassine Bechqito , Mohamed Adel Serhani , Elarbi Badidi , Jamal Riffi , Hamid Tairi

Learning domain-invariant representations has become a popular approach to unsupervised domain adaptation and is often justified by invoking a particular suite of theoretical results. We argue that there are two significant flaws in such…

Machine Learning · Statistics 2019-07-05 Fredrik D. Johansson , David Sontag , Rajesh Ranganath

We investigate smooth approximations of functions, with prescribed gradient behavior on a distinguished stratified subset of the domain. As an application, we outline how our results yield important consequences for a recently introduced…

Classical Analysis and ODEs · Mathematics 2015-07-21 D. Drusvyatskiy , M. Larsson
‹ Prev 1 8 9 10 Next ›