中文
相关论文

相关论文: The convergence of the Generalized Lanczos Trust-R…

200 篇论文

We propose a trust-region type method for a class of nonsmooth nonconvex optimization problems where the objective function is a summation of a (probably nonconvex) smooth function and a (probably nonsmooth) convex function. The model…

最优化与控制 · 数学 2021-10-26 Ziang Chen , Andre Milzarek , Zaiwen Wen

We consider the minimization of a cost function $f$ on a manifold $M$ using Riemannian gradient descent and Riemannian trust regions (RTR). We focus on satisfying necessary optimality conditions within a tolerance $\varepsilon$.…

最优化与控制 · 数学 2018-05-01 Nicolas Boumal , P. -A. Absil , Coralia Cartis

GMRES is a popular Krylov subspace method for solving linear systems of equations involving a general non-Hermitian coefficient matrix. The conventional bounds on GMRES convergence involve polynomial approximation problems in the complex…

数值分析 · 数学 2022-09-07 Mark Embree

In this paper, we propose and analyze a trust-region model-based algorithm for solving unconstrained stochastic optimization problems. Our framework utilizes random models of an objective function $f(x)$, obtained from stochastic…

最优化与控制 · 数学 2016-09-26 Ruobing Chen , Matt Menickelly , Katya Scheinberg

Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollout policy $\pi_{\text{roll}}$. However, modern LLM-RL pipelines suffer from unavoidable…

机器学习 · 计算机科学 2026-03-02 Yingru Li , Jiacai Liu , Jiawei Xu , Yuxuan Tong , Ziniu Li , Qian Liu , Baoxiang Wang

This paper proposes a random subspace trust-region algorithm for general convex-constrained derivative-free optimization (DFO) problems. Similar to previous random subspace DFO methods, the convergence of our algorithm requires a certain…

最优化与控制 · 数学 2026-05-14 Yiwen Chen , Warren Hare , Amy Wiebe

In this paper, we study second-order algorithms for solving nonconvex-strongly concave minimax problems, which have attracted much attention in recent years in many fields, especially in machine learning.We propose a gradient norm…

最优化与控制 · 数学 2025-06-17 Jun-Lin Wang , Zi Xu

We describe a Lanczos-based algorithm for approximating the product of a rational matrix function with a vector. This algorithm, which we call the Lanczos method for optimal rational matrix function approximation (Lanczos-OR), returns the…

数值分析 · 数学 2023-06-01 Tyler Chen , Anne Greenbaum , Cameron Musco , Christopher Musco

Trust region policy optimization (TRPO) is a popular and empirically successful policy search algorithm in Reinforcement Learning (RL) in which a surrogate problem, that restricts consecutive policies to be 'close' to one another, is…

机器学习 · 计算机科学 2019-12-13 Lior Shani , Yonathan Efroni , Shie Mannor

We develop a trust-region method for minimizing the sum of a smooth term $f$ and a nonsmooth term $h$), both of which can be nonconvex. Each iteration of our method minimizes a possibly nonconvex model of $f + h$ in a trust region. The…

最优化与控制 · 数学 2021-08-04 Aleksandr Y. Aravkin , Robert Baraldi , Dominique Orban

A stochastic second-order trust region method is proposed, which can be viewed as a second-order extension of the trust-region-ish (TRish) algorithm proposed by Curtis et al. (INFORMS J. Optim. 1(3) 200-220, 2019). In each iteration, a…

最优化与控制 · 数学 2019-11-19 Frank E. Curtis , Rui Shi

GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to brittle optimization behavior, as global importance-ratio…

机器学习 · 计算机科学 2026-02-09 Doyeon Lee , Eunyi Lyou , Hyunsoo Cho , Sookyung Kim , Joonseok Lee , Jaemoo Choi

We target the problem of finding a local minimum in non-convex finite-sum minimization. Towards this goal, we first prove that the trust region method with inexact gradient and Hessian estimation can achieve a convergence rate of order…

最优化与控制 · 数学 2019-03-06 Zebang Shen , Pan Zhou , Cong Fang , Alejandro Ribeiro

The residual cutting (RC) method has been proposed for efficiently solving linear equations obtained from elliptic partial differential equations. Based on the RC, we have introduced the generalized residual cutting (GRC) method, which can…

数值分析 · 计算机科学 2018-02-02 Toshihiko Abe , Anthony Theodore Chronopoulos

Currently, existing tensor recovery methods fail to recognize the impact of tensor scale variations on their structural characteristics. Furthermore, existing studies face prohibitive computational costs when dealing with large-scale…

机器学习 · 计算机科学 2025-07-09 Wenjin Qin , Hailin Wang , Jingyao Hou , Jianjun Wang

We introduce two multifidelity trust-region methods based on the Magical Trust Region (MTR) framework. MTR augments the classical trust-region step with a secondary, informative direction. In our approaches, the secondary ``magical''…

In recent years, random subspace methods have been actively studied for large-dimensional nonconvex problems. Recent subspace methods have improved theoretical guarantees such as iteration complexity and local convergence rate while…

最优化与控制 · 数学 2025-03-25 Rei Higuchi , Pierre-Louis Poirion , Akiko Takeda

We propose a stochastic first-order trust-region method with inexact function and gradient evaluations for solving finite-sum minimization problems. Using a suitable reformulation of the given problem, our method combines the inexact…

最优化与控制 · 数学 2022-10-25 Stefania Bellavia , Natasa Krejic , Benedetta Morini , Simone Rebegoldi

Trust-region (TR) type method, based on a quadratic model such as the trust-region subproblem (TRS) and $ p $-regularization subproblem ($p$RS), is arguably one of the most successful methods for unconstrained minimization. In this paper,…

最优化与控制 · 数学 2021-09-07 Liaoyuan Zeng , Ting Kei Pong

For optimization problems with nonlinear constraints, linearly constrained Lagrangian (LCL) methods sequentially minimize a Lagrangian function subject to linearized constraints. These methods converge rapidly near a solution but may not be…

最优化与控制 · 数学 2007-05-23 Michael P. Friedlander , Michael A Saunders