English
Related papers

Related papers: Descent modulus and applications

200 papers

We introduce $\mathbf{G}$radient Descent with $\mathbf{A}$daptive $\mathbf{M}$omentum $\mathbf{S}$caling ($\mathbf{Grams}$), a novel optimization algorithm that decouples the direction and magnitude of parameter updates in deep learning.…

Machine Learning · Computer Science 2025-03-06 Yang Cao , Xiaoyu Li , Zhao Song

When smoothing a function $f$ via convolution with some kernel, it is often desirable to adapt the amount of smoothing locally to the variation of $f$. For this purpose, the constant smoothing coefficient of regular convolutions needs to be…

Functional Analysis · Mathematics 2018-05-08 Ilja Klebanov

An influential line of recent work has focused on the generalization properties of unregularized gradient-based learning procedures applied to separable linear classification with exponentially-tailed loss functions. The ability of such…

Machine Learning · Computer Science 2022-06-24 Matan Schliserman , Tomer Koren

We consider the problem of maximizing non-negative non-decreasing set functions. Although most of the recent work focus on exploiting submodularity, it turns out that several objectives we encounter in practice are not submodular.…

Data Structures and Algorithms · Computer Science 2018-06-19 Gaurav Gupta , Sergio Pequito , Paul Bogdan

Consider a class of functions of one real variable with the following uniqueness property: if a function f(x) from the class vanishes on a set of positive measure, then f is the zero function. In many instances, we would like to have a…

Classical Analysis and ODEs · Mathematics 2007-05-23 F. Nazarov , M. Sodin , A. Volberg

Many statistical $M$-estimators are based on convex optimization problems formed by the combination of a data-dependent loss function with a norm-based regularizer. We analyze the convergence rates of projected gradient and composite…

Machine Learning · Statistics 2012-07-26 Alekh Agarwal , Sahand N. Negahban , Martin J. Wainwright

Shadowing trajectories are model trajectories consistent with a sequence of observations of a system, given a distribution of observational noise. The existence of such trajectories is a desirable property of any forecast model. Gradient…

Data Analysis, Statistics and Probability · Physics 2019-09-16 Roland M. B. Young , Roman Binter , Falk Niehörster , Peter L. Read , Leonard A. Smith

Many tasks in machine learning and signal processing can be solved by minimizing a convex function of a measure. This includes sparse spikes deconvolution or training a neural network with a single hidden layer. For these problems, we study…

Optimization and Control · Mathematics 2018-10-30 Lenaic Chizat , Francis Bach

In deep learning, classification tasks are formalized as optimization problems often solved via the minimization of the cross-entropy. However, recent advancements in the design of objective functions allow the usage of the $f$-divergence…

Machine Learning · Computer Science 2024-05-17 Nicola Novello , Andrea M. Tonello

In this paper, we address the optimization problem of minimizing $Q(df_x)$ over a Hadamard manifold ${\cal M}$, where $f$ is a convex function on ${\cal M}$, $df_x$ is the differential of $f$ at $x \in {\cal M}$, and $Q$ is a function on…

Optimization and Control · Mathematics 2026-01-23 Hiroshi Hirai

A generalization of the classical Sard theorem in the plane is the following. Let $f$ be a function defined on a subset $A\subset{\mathbb R}^2$. If $f$ has modulus of continuity $\omega(r)\lesssim r^2$, then $f(A)\subset{\mathbb R}$ has…

Classical Analysis and ODEs · Mathematics 2025-04-10 Iqra Altaf , Marianna Csörnyei

We study the iteration complexity of stochastic gradient descent (SGD) for minimizing the gradient norm of smooth, possibly nonconvex functions. We provide several results, implying that the $\mathcal{O}(\epsilon^{-4})$ upper bound of…

Machine Learning · Computer Science 2021-07-30 Yoel Drori , Ohad Shamir

We consider the gradient method with variable step size for minimizing functions that are definable in o-minimal structures on the real field and differentiable with locally Lipschitz gradients. We prove that global convergence holds if…

Optimization and Control · Mathematics 2024-12-02 Cédric Josz

We study the type of solutions to which stochastic gradient descent converges when used to train a single hidden-layer multivariate ReLU network with the quadratic loss. Our results are based on a dynamical stability analysis. In the…

Machine Learning · Computer Science 2023-07-03 Mor Shpigel Nacson , Rotem Mulayoff , Greg Ongie , Tomer Michaeli , Daniel Soudry

Stochastic Gradient Descent (SGD) based methods have been widely used for training large-scale machine learning models that also generalize well in practice. Several explanations have been offered for this generalization performance, a…

Machine Learning · Computer Science 2021-02-11 Yikai Zhang , Wenjia Zhang , Sammy Bald , Vamsi Pingali , Chao Chen , Mayank Goswami

In this paper, a sequential search method for finding the global minimum of an objective function is presented, The descent gradient search is repeated until the global minimum is obtained. The global minimum is located by a process of…

Optimization and Control · Mathematics 2024-02-06 Mohamed Tifroute , Anouar Lahmdani , Hassane Bouzahir

We define a number of natural (from geometric and combinatorial points of view) deformation spaces of valuations on finite graphs, and study functions over these deformation spaces. These functions include both direct metric invariants…

Combinatorics · Mathematics 2007-05-23 Dmitry Jakobson , Igor Rivin

We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case constants. Key to our proofs is directional smoothness, a…

Machine Learning · Computer Science 2025-01-15 Aaron Mishkin , Ahmed Khaled , Yuanhao Wang , Aaron Defazio , Robert M. Gower

In many naturally occurring optimization problems one needs to ensure that the definition of the optimization problem lends itself to solutions that are tractable to compute. In cases where exact solutions cannot be computed tractably, it…

Machine Learning · Computer Science 2015-05-08 Bharath Sankaran , Marjan Ghazvininejad , Xinran He , David Kale , Liron Cohen

Convolutional neural networks are widely used in imaging and image recognition. Learning such networks from training data leads to the minimization of a non-convex function. This makes the analysis of standard optimization methods such as…

Optimization and Control · Mathematics 2026-01-14 Jona-Maria Diederen , Holger Rauhut , Ulrich Terstiege
‹ Prev 1 8 9 10 Next ›