中文
相关论文

相关论文: Primal-dual regression approach for Markov decisio…

200 篇论文

This paper develops algorithms for high-dimensional stochastic control problems based on deep learning and dynamic programming. Unlike classical approximate dynamic programming approaches, we first approximate the optimal policy by means of…

概率论 · 数学 2021-09-21 Côme Huré , Huyên Pham , Achref Bachouch , Nicolas Langrené

We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove…

机器学习 · 计算机科学 2011-05-02 Shie Mannor , John Tsitsiklis

We propose and study a novel stochastic inertial primal-dual approach to solve composite optimization problems. These latter problems arise naturally when learning with penalized regularization schemes. Our analysis provide convergence…

最优化与控制 · 数学 2015-07-06 Lorenzo Rosasco , Silvia Villa , Bang Cong Vu

This article develops a primal dual formulation for a primal proximal approach suitable for a large class of non-convex models in the calculus of variations. The results are established through standard tools of functional analysis, convex…

最优化与控制 · 数学 2021-07-27 Fabio Silva Botelho

We study Concave Constrained Markov Decision Processes (Concave CMDPs) where both the objective and constraints are defined as concave functions of the state-action occupancy measure. We propose the Variance-Reduced Primal-Dual Policy…

机器学习 · 计算机科学 2024-05-28 Donghao Ying , Mengzi Amy Guo , Hyunin Lee , Yuhao Ding , Javad Lavaei , Zuo-Jun Max Shen

We present metrics for measuring state similarity in Markov decision processes (MDPs) with infinitely many states, including MDPs with continuous state spaces. Such metrics provide a stable quantitative analogue of the notion of…

人工智能 · 计算机科学 2012-07-09 Norman Ferns , Prakash Panangaden , Doina Precup

In this paper we propose and analyze two dual methods based on inexact gradient information and averaging that generate approximate primal solutions for smooth convex optimization problems. The complicating constraints are moved into the…

最优化与控制 · 数学 2013-02-14 Ion Necoara , Valentin Nedelcu

We propose an algorithm-independent framework to equip existing optimization methods with primal-dual certificates. Such certificates and corresponding rate of convergence guarantees are important for practitioners to diagnose progress, in…

机器学习 · 计算机科学 2016-06-06 Celestine Dünner , Simone Forte , Martin Takáč , Martin Jaggi

We study a stochastic primal-dual method for constrained optimization over Riemannian manifolds with bounded sectional curvature. We prove non-asymptotic convergence to the optimal objective value. More precisely, for the class of…

最优化与控制 · 数学 2017-03-24 Masoud Badiei Khuzani , Na Li

In this paper, we study a mean-variance optimization problem in an infinite horizon discrete time discounted Markov decision process (MDP). The objective is to minimize the variance of system rewards with the constraint of mean performance.…

最优化与控制 · 数学 2017-08-24 Li Xia

This paper concerns discrete-time infinite-horizon stochastic control systems with Borel state and action spaces and universally measurable policies. We study optimization problems on strategic measures induced by the policies in these…

最优化与控制 · 数学 2023-12-22 Huizhen Yu

We develop a novel primal-dual algorithm to solve a class of nonsmooth and nonlinear compositional convex minimization problems, which covers many existing and brand-new models as special cases. Our approach relies on a combination of a new…

最优化与控制 · 数学 2021-04-20 Yuzixuan Zhu , Deyi Liu , Quoc Tran-Dinh

In this note, we provide an overarching analysis of primal-dual dynamics associated to linear equality-constrained optimization problems using contraction analysis. For the well-known standard version of the problem: we establish…

系统与控制 · 电气工程与系统科学 2021-06-22 Pedro Cisneros-Velarde , Saber Jafarpour , Francesco Bullo

We propose a doubly stochastic primal-dual coordinate optimization algorithm for empirical risk minimization, which can be formulated as a bilinear saddle-point problem. In each iteration, our method randomly samples a block of coordinates…

机器学习 · 计算机科学 2017-04-13 Adams Wei Yu , Qihang Lin , Tianbao Yang

We consider policy evaluation in infinite-horizon discounted Markov decision problems (MDPs) with infinite spaces. We reformulate this task a compositional stochastic program with a function-valued decision variable that belongs to a…

最优化与控制 · 数学 2020-05-19 Alec Koppel , Garrett Warnell , Ethan Stump , Peter Stone , Alejandro Ribeiro

This paper develops a distributed model predictive control (DMPC) strategy for a class of discrete-time linear systems with consideration of globally coupled constraints. The DMPC under study is based on the dual problem concerning all…

最优化与控制 · 数学 2019-07-25 Yanxu Su , Yang Shi , Changyin Sun

We study the minimization of a spectral risk measure of the total discounted cost generated by a Markov Decision Process (MDP) over a finite or infinite planning horizon. The MDP is assumed to have Borel state and action spaces and the cost…

最优化与控制 · 数学 2025-10-16 Nicole Bäuerle , Alexander Glauner

Presented is a method for efficient computation of the Hamilton-Jacobi (HJ) equation for time-optimal control problems using the generalized Hopf formula. Typically, numerical methods to solve the HJ equation rely on a discrete grid of the…

系统与控制 · 计算机科学 2019-10-22 Matthew R. Kirchner , Gary Hewer , Jerome Darbon , Stanley Osher

The curse of dimensionality is a widely known issue in reinforcement learning (RL). In the tabular setting where the state space $\mathcal{S}$ and the action space $\mathcal{A}$ are both finite, to obtain a nearly optimal policy with…

机器学习 · 计算机科学 2022-10-28 Bingyan Wang , Yuling Yan , Jianqing Fan

The paper addresses two variants of the stochastic shortest path problem ('optimize the accumulated weight until reaching a goal state') in Markov decision processes (MDPs) with integer weights. The first variant optimizes partial expected…

计算机科学中的逻辑 · 计算机科学 2019-05-01 Jakob Piribauer , Christel Baier
‹ 上一页 1 8 9 10 下一页 ›