English
Related papers

Related papers: Stochastic optimization on matrices and a graphon …

200 papers

We present a new approach, based on graphon theory, to finding the limiting spectral distributions of general Wigner-type matrices. This approach determines the moments of the limiting measures and the equations of their Stieltjes…

Probability · Mathematics 2020-08-11 Yizhe Zhu

The stochastic gradient Langevin Dynamics is one of the most fundamental algorithms to solve sampling problems and non-convex optimization appearing in several machine learning applications. Especially, its variance reduced versions have…

Machine Learning · Computer Science 2022-11-22 Yuri Kinoshita , Taiji Suzuki

We study the implicit regularization of mini-batch stochastic gradient descent, when applied to the fundamental problem of least squares regression. We leverage a continuous-time stochastic differential equation having the same moments as…

Machine Learning · Statistics 2020-06-23 Alnur Ali , Edgar Dobriban , Ryan J. Tibshirani

We develop a new continuous-time stochastic gradient descent method for optimizing over the stationary distribution of stochastic differential equation (SDE) models. The algorithm continuously updates the SDE model's parameters using an…

Machine Learning · Computer Science 2023-08-29 Ziheng Wang , Justin Sirignano

Symmetries are prevalent in deep learning and can significantly influence the learning dynamics of neural networks. In this paper, we examine how exponential symmetries -- a broad subclass of continuous symmetries present in the model…

Machine Learning · Computer Science 2024-11-08 Liu Ziyin , Mingze Wang , Hongchao Li , Lei Wu

Stochastic gradient descent (SGD) is widely used in deep learning due to its computational efficiency, but a complete understanding of why SGD performs so well remains a major challenge. It has been observed empirically that most…

Machine Learning · Statistics 2022-06-20 Carmina Fjellström , Kaj Nyström

Graph sparsification has been studied extensively over the past two decades, culminating in spectral sparsifiers of optimal size (up to constant factors). Spectral hypergraph sparsification is a natural analogue of this problem, for which…

Data Structures and Algorithms · Computer Science 2021-06-07 Michael Kapralov , Robert Krauthgamer , Jakab Tardos , Yuichi Yoshida

We study graphons as a non-parametric generalization of stochastic block models, and show how to obtain compactly represented estimators for sparse networks in this framework. Our algorithms and analysis go beyond previous work in several…

Statistics Theory · Mathematics 2016-02-25 Christian Borgs , Jennifer T. Chayes , Henry Cohn , Shirshendu Ganguly

The paper deals with the convergence properties of the products of random (row-)stochastic matrices. The limiting behavior of such products is studied from a dynamical system point of view. In particular, by appropriately defining a dynamic…

Probability · Mathematics 2013-01-15 Behrouz Touri , Angelia Nedich

Stochastic optimization algorithms update models with cheap per-iteration costs sequentially, which makes them amenable for large-scale data analysis. Such algorithms have been widely studied for structured sparse models where the sparsity…

Machine Learning · Computer Science 2019-05-10 Baojian Zhou , Feng Chen , Yiming Ying

This work establishes rigorous, novel and widely applicable stability guarantees and transferability bounds for graph convolutional networks -- without reference to any underlying limit object or statistical distribution. Crucially,…

Machine Learning · Computer Science 2023-10-03 Christian Koke

In this paper, we exploit the theory of dense graph limits to provide a new framework to study the stability of graph partitioning methods, which we call structural consistency. Both stability under perturbation as well as asymptotic…

Combinatorics · Mathematics 2016-08-15 Peter Diao , Dominique Guillot , Apoorva Khare , Bala Rajaratnam

We consider a countable system of interacting (possibly non-Markovian) stochastic differential equations driven by independent Brownian motions and indexed by the vertices of a locally finite graph $G = (V,E)$. The drift of the process at…

Probability · Mathematics 2020-09-28 Daniel Lacker , Kavita Ramanan , Ruoyu Wu

We study Turing bifurcations on one-dimensional random ring networks where the probability of a connection between two nodes depends on the distance between the two nodes. Our approach uses the theory of graphons to approximate the graph…

Dynamical Systems · Mathematics 2026-03-03 Jason Bramburger , Matt Holzer

Stochastic gradient descent (SGD) is a popular algorithm for minimizing objective functions that arise in machine learning. For constant step-sized SGD, the iterates form a Markov chain on a general state space. Focusing on a class of…

Optimization and Control · Mathematics 2025-03-26 David Shirokoff , Philip Zaleski

Two frameworks that have been used to characterize reflected diffusions include stochastic differential equations with reflection and the so-called submartingale problem. We introduce a general formulation of the submartingale problem for…

Probability · Mathematics 2014-12-03 Weining Kang , Kavita Ramanan

In this paper, we show a large deviation principle for certain sequences of static Schr\"{o}dinger bridges, typically motivated by a scale-parameter decreasing towards zero, extending existing large deviation results to cover a wider range…

Probability · Mathematics 2025-06-23 Viktor Nilsson , Pierre Nyquist

We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps (CFMs): dynamical systems that evolve a distinguished token in the presence of a contextual measure…

Machine Learning · Computer Science 2026-05-19 Shi Chen , Zhengjiang Lin , Kaizhao Liu , Philippe Rigollet

We design algorithms for fitting a high-dimensional statistical model to a large, sparse network without revealing sensitive information of individual members. Given a sparse input graph $G$, our algorithms output a…

Statistics Theory · Mathematics 2015-06-23 Christian Borgs , Jennifer T. Chayes , Adam Smith

The optimistic gradient method is useful in addressing minimax optimization problems. Motivated by the observation that the conventional stochastic version suffers from the need for a large batch size on the order of…

Machine Learning · Computer Science 2024-01-29 Haoyuan Cai , Sulaiman A. Alghunaim , Ali H. Sayed