English
Related papers

Related papers: Local Optimality and Generalization Guarantees for…

200 papers

This paper introduces a new method to train recurrent neural networks using dynamical trajectory-based optimization. The optimization method utilizes a projected gradient system (PGS) and a quotient gradient system (QGS) to determine the…

Signal Processing · Electrical Eng. & Systems 2019-10-16 Hamid Khodabandehlou , M. Sami Fadali

We study Langevin-type algorithms for sampling from Gibbs distributions such that the potentials are dissipative and their weak gradients have finite moduli of continuity not necessarily convergent to zero. Our main result is a…

Statistics Theory · Mathematics 2024-03-01 Shogo Nakakita

The Expectation-Maximization (EM) algorithm is a popular choice for learning latent variable models. Variants of the EM have been initially introduced, using incremental updates to scale to large datasets, and using Monte Carlo (MC)…

Machine Learning · Statistics 2022-03-22 Belhal Karimi , Ping Li

In this paper, we study the Empirical Risk Minimization problem in the non-interactive local model of differential privacy. In the case of constant or low dimensionality ($p\ll n$), we first show that if the ERM loss function is $(\infty,…

Machine Learning · Computer Science 2018-05-18 Di Wang , Marco Gaboardi , Jinhui Xu

In the context of first-order algorithms subject to random gradient noise, we study the trade-offs between the convergence rate (which quantifies how fast the initial conditions are forgotten) and the "risk" of suboptimality, i.e.…

Optimization and Control · Mathematics 2025-03-11 Bugra Can , Mert Gürbüzbalaban

We propose a new method called the Metropolis-adjusted Mirror Langevin algorithm for approximate sampling from distributions whose support is a compact and convex set. This algorithm adds an accept-reject filter to the Markov chain induced…

Computation · Statistics 2024-06-24 Vishwak Srinivasan , Andre Wibisono , Ashia Wilson

The Metropolis-adjusted Langevin (MALA) algorithm is a sampling algorithm which makes local moves by incorporating information about the gradient of the logarithm of the target density. In this paper we study the efficiency of MALA on a…

Probability · Mathematics 2012-11-29 Natesh S. Pillai , Andrew M. Stuart , Alexandre H. Thiéry

The Levenberg-Marquardt algorithm is a flexible iterative procedure used to solve non-linear least squares problems. In this work we study how a class of possible adaptations of this procedure can be used to solve maximum likelihood…

Computation · Statistics 2014-10-06 Marco Giordan , Federico Vaggi , Ron Wehrens

In this paper, we investigate a continuous time version of the Stochastic Langevin Monte Carlo method, introduced in [WT11], that incorporates a stochastic sampling step inside the traditional over-damped Langevin diffusion. This method is…

Machine Learning · Statistics 2023-01-10 Marelys Crespo Navas , Sébastien Gadat , Xavier Gendre

Robotic systems must be able to quickly and robustly make decisions when operating in uncertain and dynamic environments. While Reinforcement Learning (RL) can be used to compute optimal policies with little prior knowledge about the…

Robotics · Computer Science 2016-09-13 Yunpeng Pan , Xinyan Yan , Evangelos Theodorou , Byron Boots

Recent analyses of certain gradient descent optimization methods have shown that performance can degrade in some settings - such as with stochasticity or implicit momentum. In deep reinforcement learning (Deep RL), such optimization methods…

Machine Learning · Computer Science 2018-10-08 Peter Henderson , Joshua Romoff , Joelle Pineau

Inverse optimization involves inferring unknown parameters of an optimization problem from known solutions and is widely used in fields such as transportation, power systems, and healthcare. We study the contextual inverse optimization…

Machine Learning · Computer Science 2024-06-06 Saurabh Mishra , Anant Raj , Sharan Vaswani

In this paper, we study the generalization performance of global minima for implementing empirical risk minimization (ERM) on over-parameterized deep ReLU nets. Using a novel deepening scheme for deep ReLU nets, we rigorously prove that…

Machine Learning · Computer Science 2023-03-01 Shao-Bo Lin , Yao Wang , Ding-Xuan Zhou

We revisit the well-studied problem of differentially private empirical risk minimization (ERM). We show that for unconstrained convex generalized linear models (GLMs), one can obtain an excess empirical risk of $\tilde…

Cryptography and Security · Computer Science 2021-03-04 Shuang Song , Thomas Steinke , Om Thakkar , Abhradeep Thakurta

We study the minimal error of the Empirical Risk Minimization (ERM) procedure in the task of regression, both in the random and the fixed design settings. Our sharp lower bounds shed light on the possibility (or impossibility) of adapting…

Statistics Theory · Mathematics 2021-02-25 Gil Kur , Alexander Rakhlin

We consider the inverse Ising problem, i.e. the inference of network couplings from observed spin trajectories for a model with continuous time Glauber dynamics. By introducing two sets of auxiliary latent random variables we render the…

Machine Learning · Statistics 2017-12-22 Christian Donner , Manfred Opper

We explore past and recent developments in rare-event probability estimation with a particular focus on a novel Monte Carlo technique Empirical Likelihood Maximization (ELM). This is a versatile method that involves sampling from a sequence…

Computation · Statistics 2013-12-12 A. Huang , Z. I. Botev

The classical Langevin Monte Carlo method looks for samples from a target distribution by descending the samples along the gradient of the target distribution. The method enjoys a fast convergence rate. However, the numerical cost is…

Machine Learning · Statistics 2025-03-07 Zhiyan Ding , Qin Li

In average reward Markov decision processes, state-of-the-art algorithms for regret minimization follow a well-established framework: They are model-based, optimistic and episodic. First, they maintain a confidence region from which…

Machine Learning · Computer Science 2025-12-01 Victor Boone , Bruno Gaujal

We study the hitting times of Markov processes to target set $G$, starting from a reference configuration $x_0$ or its basin of attraction. The configuration $x_0$ can correspond to the bottom of a (meta)stable well, while the target $G$…

Probability · Mathematics 2014-06-11 R. Fernandez , F. Manzo , F. R. Nardi , E. Scoppola