English
Related papers

Related papers: Two-Point Deterministic Equivalence for Stochastic…

200 papers

In neural networks with binary activations and or binary weights the training by gradient descent is complicated as the model has piecewise constant response. We consider stochastic binary networks, obtained by adding noises in front of…

Machine Learning · Statistics 2020-11-05 Alexander Shekhovtsov , Viktor Yanush , Boris Flach

We prove that all 'gradient span algorithms' have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to infinity. In particular, this result explains the counterintuitive phenomenon that…

Machine Learning · Statistics 2024-10-15 Felix Benning , Leif Döring

Backpropagation with gradient descent is a common optimization strategy employed by most neural network architectures in machine learning. However, finding optimal hyperparameters to guide training has proven challenging. While it is widely…

Machine Learning · Computer Science 2026-05-20 Vy Bui , Hang Yu , Karthik Kantipudi , Ziv Yaniv , Stefan Jaeger

Randomness is ubiquitous in modern engineering. The uncertainty is often modeled as random coefficients in the differential equations that describe the underlying physics. In this work, we describe a two-step framework for numerically…

Numerical Analysis · Mathematics 2021-02-03 Ting Wang , Jaroslaw Knap

The exact solution for a system with two-particle annihilation and decoagulation has been studied. The spectrum of the Hamiltonian of the system is found. It is shown that the steady state is two-fold degenerate. The average number density…

Condensed Matter · Physics 2012-07-27 Amir Aghamohammadi , Mohammad Khorrami

In this paper, we present some theoretical work to explain why simple gradient descent methods are so successful in solving non-convex optimization problems in learning large-scale neural networks (NN). After introducing a mathematical tool…

Machine Learning · Computer Science 2023-05-01 Hui Jiang

The root-cause diagnostics of product quality defects in multistage manufacturing processes often requires a joint identification of crucial stages and process variables. To meet this requirement, this paper proposes a novel penalized…

Applications · Statistics 2020-06-11 Cheoljoon Jeong , Xiaolei Fang

A determinantal approximation is obtained for the permanent of a doubly stochastic matrix. For moderate-deviation matrix sequences, the asymptotic relative error is of order $O(n^{-1})$.

Combinatorics · Mathematics 2012-05-28 Peter McCullagh

We study Fredholm determinants of a class of integral operators, whose kernels can be expressed as double contour integrals of a special type. Such Fredholm determinants appear in various random matrix and statistical physics models. We…

Mathematical Physics · Physics 2020-10-29 Mattia Cafasso , Tom Claeys , Manuela Girotti

Stochastic gradient methods enable learning probabilistic models from large amounts of data. While large step-sizes (learning rates) have shown to be best for least-squares (e.g., Gaussian noise) once combined with parameter averaging,…

Machine Learning · Statistics 2018-11-22 Dmitry Babichev , Francis Bach

In this paper we discuss some properties of resolvents of an accretive operator in linear 2-normed spaces, focusing on the concept of contraction mapping and the unique fixed point of contraction mappings in linear 2- normed spaces. Also,…

Functional Analysis · Mathematics 2017-07-11 P. K. Harikrishnan , K. T. Ravindran

In a real Hilbert space setting, we study the convergence properties of an inexact gradient algorithm featuring both viscous and Hessian driven damping for convex differentiable optimization. In this algorithm, the gradient evaluation can…

Optimization and Control · Mathematics 2025-09-25 Harsh Choudhary , Jalal Fadili , Vyachelav Kungurtsev

Stochastic coordinate descent algorithms are efficient methods in which each iterate is obtained by fixing most coordinates at their values from the current iteration, and approximately minimizing the objective with respect to the remaining…

Machine Learning · Statistics 2025-04-02 Eméric Gbaguidi

Double descent refers to the phase transition that is exhibited by the generalization error of unregularized learning models when varying the ratio between the number of parameters and the number of training samples. The recent success of…

Machine Learning · Computer Science 2020-06-19 Michał Dereziński , Feynman Liang , Michael W. Mahoney

We introduce a doubly stochastic marked point process model for supervised classification problems. Regardless of the number of classes or the dimension of the feature space, the model requires only 2--3 parameters for the covariance…

Methodology · Statistics 2012-07-20 Jie Yang , Klaus Miescke , Peter McCullagh

We propose a nonconvex estimator for joint multivariate regression and precision matrix estimation in the high dimensional regime, under sparsity constraints. A gradient descent algorithm with hard thresholding is developed to solve the…

Machine Learning · Statistics 2016-06-03 Jinghui Chen , Quanquan Gu

This paper is part of a program to understand the parameter spaces of dynamical systems generated by meromorphic functions with finitely many singular values. We give a full description of the parameter space for a specific family based on…

Complex Variables · Mathematics 2022-01-24 Tao Chen , Yunping Jiang , Linda Keen

We propose a stochastic extension of the primal-dual hybrid gradient algorithm studied by Chambolle and Pock in 2011 to solve saddle point problems that are separable in the dual variable. The analysis is carried out for general…

Optimization and Control · Mathematics 2018-04-11 Antonin Chambolle , Matthias J. Ehrhardt , Peter Richtárik , Carola-Bibiane Schönlieb

In a variety of problems originating in supervised, unsupervised, and reinforcement learning, the loss function is defined by an expectation over a collection of random variables, which might be part of a probabilistic model or the external…

Machine Learning · Computer Science 2016-01-06 John Schulman , Nicolas Heess , Theophane Weber , Pieter Abbeel

An effective method for generating linear equations of maximal symmetry in their much general normal form is obtained. In the said normal form, the coefficients of the equation are differential functions of the coefficient of the term of…

Classical Analysis and ODEs · Mathematics 2015-02-26 JC Ndogmo