Related papers: Worst Exponential Decay Rate for Degenerate Gradie…
The monotone variational inequality is a central problem in mathematical programming that unifies and generalizes many important settings such as smooth convex optimization, two-player zero-sum games, convex-concave saddle point problems,…
Deep learning experiments by Cohen et al. [2021] using deterministic Gradient Descent (GD) revealed an Edge of Stability (EoS) phase when learning rate (LR) and sharpness (i.e., the largest eigenvalue of Hessian) no longer behave as in…
We systematically develop a learning-based treatment of stochastic optimal control (SOC), relying on direct optimization of parametric control policies. We propose a derivation of adjoint sensitivity results for stochastic differential…
In this work, the multiplier method is extended to obtain a general lower bound of the exponential decay rate in terms of the physical parameters for port-Hamiltonian systems in one space dimension with boundary dissipation. The physical…
In this paper we analyze the behaviour of the stochastic gradient descent (SGD), a widely used method in supervised learning for optimizing neural network weights via a minimization of non-convex loss functions. Since the pioneering work of…
We consider the one-dimensional degenerate parabolic equation $$ u_t - (x^\alpha u_x)_x =0 \qquad x\in(0,1),\ t \in (0,T) ,$$ controlled by a boundary force acting at the degeneracy point $x=0$. First we study the reachable targets at some…
This work investigates the global exponential stabilization of a degenerate Euler-Bernoulli beam subjected to a non uniform axial force and a delayed feedback control. First, we address the well-posedness of the system by constructing an…
This paper presents a discrete-time passivity-based analysis of the gradient descent method for a class of functions with sector-bounded gradients. Using a loop transformation, it is shown that the gradient descent method can be interpreted…
In this paper, we consider an expanding construction of a distributed control system, which is obtained by adding a new subsystem one after the other, until all $n$ subsystems, where $n \ge 2$, are included in the distributed control…
This paper shows that the implicit bias of gradient descent on linearly separable data is exactly characterized by the optimal solution of a dual optimization problem given by a smoothed margin, even for general losses. This is in contrast…
We propose a new method for generating realistic datasets with distribution shifts using any decoder-based generative model. Our approach systematically creates datasets with varying intensities of distribution shifts, facilitating a…
We propose new limiting dynamics for stochastic gradient descent in the small learning rate regime called stochastic modified flows. These SDEs are driven by a cylindrical Brownian motion and improve the so-called stochastic modified…
We address the challenge of estimating the learning rate for adaptive gradient methods used in training deep neural networks. While several learning-rate-free approaches have been proposed, they are typically tailored for steepest descent.…
We examine gradient descent on unregularized logistic regression problems, with homogeneous linear predictors on linearly separable datasets. We show the predictor converges to the direction of the max-margin (hard margin SVM) solution. The…
We study the large scale behavior of elliptic systems with stationary random coefficient that have only slowly decaying correlations. To this aim we analyze the so-called corrector equation, a degenerate elliptic equation posed in the…
This paper studies an infinite horizon optimal control problem for discrete-time linear system and quadratic criteria, both with random parameters which are independent and identically distributed with respect to time. In this general…
In this paper, we investigate the decay properties of an axisymmetric D-solutions to stationary incompressible Navier-Stokes systems in $\mathbb{R}^3$. We obtain the optimal decay rate $|{\bf u}(x)|\leq \frac{C}{|x|+1}$ for axisymmetric…
We consider the discrete memoryless degraded broadcast channels with feedback. We prove that the error probability of decoding tends to one exponentially for rates outside the capacity region and derive an explicit lower bound of this…
This paper explores the decentralized control of linear deterministic systems in which different controllers operate based on distinct state information, and extends the findings to the output feedback scenario. Assuming the controllers…
We analyze (stochastic) gradient descent (SGD) with delayed updates on smooth quasi-convex and non-convex functions and derive concise, non-asymptotic, convergence rates. We show that the rate of convergence in all cases consists of two…