中文
相关论文

相关论文: Understanding high-index saddle dynamics via numer…

200 篇论文

Stochastic gradient descent (SGD) is a cornerstone algorithm for high-dimensional optimization, renowned for its empirical successes. Recent theoretical advances have provided a deep understanding of how SGD enables feature learning in…

机器学习 · 统计学 2026-02-23 Nived Rajaraman , Yanjun Han

Stochastic gradient descent (SGD) is a standard optimization method to minimize a training error with respect to network parameters in modern neural network learning. However, it typically suffers from proliferation of saddle points in the…

机器学习 · 计算机科学 2017-11-23 Haiping Huang , Taro Toyoizumi

Numerous problems in optics, quantum physics, stability analysis, and control of dynamical systems can be brought to an optimization problem with matrix variable subjected to the symplecticity constraint. As this constraint nicely forms a…

最优化与控制 · 数学 2022-11-18 Bin Gao , Nguyen Thanh Son , Tatjana Stykel

Being able to efficiently obtain an accurate estimate of the failure probability of SRAM components has become a central issue as model circuits shrink their scale to submicrometer with advanced technology nodes. In this work, we revisit…

机器学习 · 计算机科学 2023-08-01 Yanfang Liu , Guohao Dai , Wei W. Xing

Stochastic gradient descent (SGD) is a prevalent optimization technique for large-scale distributed machine learning. While SGD computation can be efficiently divided between multiple machines, communication typically becomes a bottleneck…

机器学习 · 计算机科学 2021-05-24 Dmitrii Avdiukhin , Grigory Yaroslavtsev

In the last two decades, significant effort has been put in understanding and designing so-called structure-preserving numerical methods for the simulation of mechanical systems. Geometric integrators attempt to preserve the geometry…

数值分析 · 数学 2018-10-26 David Martín de Diego , Rodrigo T. Sato Martín de Almagro

The widely-adopted discretisation of the horizontal pressure gradient term formulated by Simmons and Burridge (1981) for atmospheric models on $\sigma$-$p$ hybrid vertical coordinate is found to incur spectral blocking for rotational wind…

大气与海洋物理 · 物理学 2019-09-02 Masashi Ujiie , Daisuke Hotta

This paper focuses on the distributed optimization of stochastic saddle point problems. The first part of the paper is devoted to lower bounds for the centralized and decentralized distributed methods for smooth (strongly) convex-(strongly)…

机器学习 · 计算机科学 2025-04-28 Aleksandr Beznosikov , Valentin Samokhin , Alexander Gasnikov

The optimization with orthogonality has been shown useful in training deep neural networks (DNNs). To impose orthogonality on DNNs, both computational efficiency and stability are important. However, existing methods utilizing Riemannian…

机器学习 · 计算机科学 2022-07-12 Fanchen Bu , Dong Eui Chang

In this paper, we present a stochastic augmented Lagrangian approach on (possibly infinite-dimensional) Riemannian manifolds to solve stochastic optimization problems with a finite number of deterministic constraints.We investigate the…

最优化与控制 · 数学 2025-04-01 Caroline Geiersbach , Tim Suchan , Kathrin Welker

This work considers optimization of composition of functions in a nested form over Riemannian manifolds where each function contains an expectation. This type of problems is gaining popularity in applications such as policy evaluation in…

最优化与控制 · 数学 2024-03-20 Dewei Zhang , Sam Davanloo Tajbakhsh

The information exponent ([BAGJ21]) and its extensions -- which are equivalent to the lowest degree in the Hermite expansion of the link function (after a potential label transform) for Gaussian single-index models -- have played an…

机器学习 · 计算机科学 2025-10-07 Yunwei Ren , Jason D. Lee

We study the properties of stochastic approximation applied to a tame nondifferentiable function subject to constraints defined by a Riemannian manifold. The objective landscape of tame functions, arising in o-minimal topology extended to a…

机器学习 · 计算机科学 2025-08-13 Johannes Aspman , Vyacheslav Kungurtsev , Reza Roohi Seraji

Policy gradient methods are an appealing approach in reinforcement learning because they directly optimize the cumulative reward and can straightforwardly be used with nonlinear function approximators such as neural networks. The two main…

机器学习 · 计算机科学 2018-10-23 John Schulman , Philipp Moritz , Sergey Levine , Michael Jordan , Pieter Abbeel

Stochastic Gradient Descent (SGD) and its Ruppert-Polyak averaged variant (ASGD) lie at the heart of modern large-scale learning, yet their theoretical properties in high-dimensional settings are rarely understood. In this paper, we provide…

机器学习 · 统计学 2025-10-15 Jiaqi Li , Zhipeng Lou , Johannes Schmidt-Hieber , Wei Biao Wu

Augmented Lagrangian (AL) methods are a well known class of algorithms for solving constrained optimization problems. They have been extended to the solution of saddle-point systems of linear equations. We study an AL (SPAL) algorithm for…

数值分析 · 数学 2024-04-24 N. Huang , Y. -H. Dai , D. Orban , M. A. Saunders

In this paper, we focus on the decentralized optimization problem over the Stiefel manifold, which is defined on a connected network of $d$ agents. The objective is an average of $d$ local functions, and each function is privately held by…

最优化与控制 · 数学 2022-07-13 Lei Wang , Xin Liu

We consider the geometric numerical integration of Hamiltonian systems subject to both equality and "hard" inequality constraints. As in the standard geometric integration setting, we target long-term structure preservation. We…

数值分析 · 数学 2011-06-02 Danny M. Kaufman , Dinesh K. Pai

Meta-learning problem is usually formulated as a bi-level optimization in which the task-specific and the meta-parameters are updated in the inner and outer loops of optimization, respectively. However, performing the optimization in the…

机器学习 · 计算机科学 2024-06-04 Hadi Tabealhojeh , Soumava Kumar Roy , Peyman Adibi , Hossein Karshenas

In industrial big data scenarios, high-dimensional sparse matrices (HDI) are widely used to characterize high-order interaction relationships among massive nodes. The stochastic gradient descent-based latent factor analysis (SGD-LFA) method…

机器学习 · 计算机科学 2025-08-26 Jinli Li , Shiyu Long , Minglian Han