中文
相关论文

相关论文: Completion vs Optimality: Policy Gradient in Long-…

200 篇论文

Online reinforcement learning in infinite-horizon Markov decision processes (MDPs) remains less theoretically and algorithmically developed than its episodic counterpart, with many algorithms suffering from high ``burn-in'' costs and…

机器学习 · 计算机科学 2026-03-26 Guy Zamir , Matthew Zurek , Yudong Chen

Policy gradient lies at the core of deep reinforcement learning (RL) in continuous domains. Despite much success, it is often observed in practice that RL training with policy gradient can fail for many reasons, even on standard control…

机器学习 · 计算机科学 2024-01-23 Tao Wang , Sylvia Herbert , Sicun Gao

In this paper, we develop a provably correct optimal control strategy for a finite deterministic transition system. By assuming that penalties with known probabilities of occurrence and dynamics can be sensed locally at the states of the…

机器人学 · 计算机科学 2013-03-15 Mária Svoreňová , Ivana Černá , Calin Belta

We consider the finite horizon continuous reinforcement learning problem. Our contribution is three-fold. First,we give a tractable algorithm based on optimistic value iteration for the problem. Next,we give a lower bound on regret of order…

机器学习 · 计算机科学 2019-08-05 Phanideep Gampa , Sairam Satwik Kondamudi , Lakshmanan Kailasam

We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real-world systems largely because it resists many concentration arguments. In this paper, we assume…

最优化与控制 · 数学 2025-11-12 Dheeraj Narasimha , Nicolas Gast

In this paper, we investigate dynamic optimization problems featuring both stochastic control and optimal stopping in a finite time horizon. The paper aims to develop new methodologies, which are significantly different from those of mixed…

投资组合管理 · 定量金融 2014-06-27 Xiongfei Jian , Xun Li , Fahuai Yi

We propose stochastic decision horizons (SDH), a theoretically grounded framework for solving constrained RL problems with every-step constraint satisfaction, a desirable property in many real-world applications. In SDH, a constraint…

机器学习 · 计算机科学 2026-05-27 Nikola Milosevic , Leonard Franz , Daniel Haeufle , Georg Martius , Nico Scherf , Pavel Kolev

Training large language model (LLM) agents for adversarial games is often driven by episodic objectives such as win rate. In long-horizon settings, however, payoffs are shaped by latent strategic externalities that evolve over time, so…

机器学习 · 计算机科学 2026-02-10 Boyang Xia , Weiyou Tian , Qingnan Ren , Jiaqi Huang , Jie Xiao , Shuo Lu , Kai Wang , Lynn Ai , Eric Yang , Bill Shi

We explore reinforcement learning methods for finding the optimal policy in the linear quadratic regulator (LQR) problem. In particular, we consider the convergence of policy gradient methods in the setting of known and unknown parameters.…

机器学习 · 计算机科学 2021-06-25 Ben Hambly , Renyuan Xu , Huining Yang

Constrained Markov Decision Processes (CMDPs) formalize sequential decision-making problems whose objective is to minimize a cost function while satisfying constraints on various cost functions. In this paper, we consider the setting of…

机器学习 · 计算机科学 2020-09-25 Krishna C. Kalagarla , Rahul Jain , Pierluigi Nuzzo

In this paper, we consider an infinite horizon average reward Markov Decision Process (MDP). Distinguishing itself from existing works within this context, our approach harnesses the power of the general policy gradient-based algorithm,…

机器学习 · 计算机科学 2024-02-06 Qinbo Bai , Washim Uddin Mondal , Vaneet Aggarwal

Optimally solving a multi-armed bandit problem suffers the curse of dimensionality. Indeed, resorting to dynamic programming leads to an exponential growth of computing time, as the number of arms and the horizon increase. We introduce a…

最优化与控制 · 数学 2024-05-22 Michel de Lara , Benjamin Heymann , Jean-Philippe Chancelier

In this work, we propose a Model Predictive Control (MPC) formulation incorporating two distinct horizons: a prediction horizon and a constraint horizon. This approach enables a deeper understanding of how constraints influence key system…

系统与控制 · 电气工程与系统科学 2025-03-25 Allan Andre Do Nascimento , Han Wang , Antonis Papachristodoulou , Kostas Margellos

A key aspect of Safe Reinforcement Learning (Safe RL) involves estimating the constraint condition for the next policy, which is crucial for guiding the optimization of safe policy updates. However, the existing Advantage-based Estimation…

机器学习 · 计算机科学 2024-12-17 Juntao Dai , Yaodong Yang , Qian Zheng , Gang Pan

We propose adaptation strategies to modify the standard constrained model predictive controller scheme in order to guarantee a certain lower bound on the degree of suboptimality. Within this analysis, the length of the optimization horizon…

最优化与控制 · 数学 2015-03-19 Jürgen Pannek

We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove…

机器学习 · 计算机科学 2011-05-02 Shie Mannor , John Tsitsiklis

In this work, solution of the finite horizon hybrid optimal control problem as the central element of the receding horizon optimal control (model predictive control) is investigated based on the indirect approach. The response of a hybrid…

系统与控制 · 计算机科学 2020-09-24 Babak Tavassoli

In this paper we get error bounds for fully discrete approximations of infinite horizon problems via the dynamic programming approach. It is well known that considering a time discretization with a positive step size $h$ an error bound of…

数值分析 · 数学 2026-02-09 Javier de Frutos , Julia Novo

Model predictive control (MPC) schemes are commonly designed with fixed, i.e., time-invariant, horizon length and cost functions. If no stabilizing terminal ingredients are used, stability can be guaranteed via a sufficiently long horizon.…

系统与控制 · 电气工程与系统科学 2021-03-02 Lukas Beckenbach , Stefan Streif

We study the problem of reinforcement learning in infinite-horizon discounted linear Markov decision processes (MDPs), and propose the first computationally efficient algorithm achieving rate-optimal regret guarantees in this setting. Our…

机器学习 · 计算机科学 2026-03-16 Antoine Moulin , Gergely Neu , Luca Viano