中文
相关论文

相关论文: Unbounded Markov Dynamic Programming with Weighted…

200 篇论文

This paper concerns computation of optimal policies in which the one-step reward function contains a cost term that models Kullback-Leibler divergence with respect to nominal dynamics. This technique was introduced by Todorov in 2007, where…

最优化与控制 · 数学 2018-07-27 Ana Bušić , Sean Meyn

With the objective of developing computational methods for stability analysis of switched systems, we consider the problem of finding the minimal lower bounds on average dwell-time that guarantee global asymptotic stability of the origin.…

最优化与控制 · 数学 2023-07-24 Sigurdur Hafstein , Aneel Tanwani

Approximate dynamic programming is a popular method for solving large Markov decision processes. This paper describes a new class of approximate dynamic programming (ADP) methods- distributionally robust ADP-that address the curse of…

机器学习 · 统计学 2012-05-22 Marek Petrik

We introduce and study constrained Markov Decision Processes (cMDPs) with anytime constraints. An anytime constraint requires the agent to never violate its budget at any point in time, almost surely. Although Markovian policies are no…

机器学习 · 计算机科学 2024-06-14 Jeremy McMahan , Xiaojin Zhu

The paper deals with finite-state Markov decision processes (MDPs) with integer weights assigned to each state-action pair. New algorithms are presented to classify end components according to their limiting behavior with respect to the…

计算机科学中的逻辑 · 计算机科学 2018-05-01 Christel Baier , Nathalie Bertrand , Clemens Dubslaff , Daniel Gburek , Ocan Sankur

Adaptive nuclear-norm penalization is proposed for low-rank matrix approximation, by which we develop a new reduced-rank estimation method for the general high-dimensional multivariate regression problems. The adaptive nuclear norm of a…

统计方法学 · 统计学 2012-09-25 Kun Chen , Hongbo Dong , Kung-Sik Chan

The problem of variable-rate lossless data compression is considered, for codes with and without prefix constraints. Sharp bounds are derived for the best achievable compression rate of memoryless sources, when the excess-rate probability…

信息论 · 计算机科学 2025-11-13 Andreas Theocharous , Lampros Gavalakis , Ioannis Kontoyiannis

Many Reinforcement Learning algorithms assume a Markov reward function to guarantee optimality. However, not all reward functions are Markov. This paper proposes a framework for mapping non-Markov reward functions into equivalent Markov…

机器学习 · 计算机科学 2024-08-19 Gregory Hyde , Eugene Santos

The design of fixed point algorithms is at the heart of monotone operator theory, convex analysis, and of many modern optimization problems arising in machine learning and control. This tutorial reviews recent advances in understanding the…

最优化与控制 · 数学 2022-07-19 Francesco Bullo , Pedro Cisneros-Velarde , Alexander Davydov , Saber Jafarpour

This paper studies the dynamic programming principle using the measurable selection method for stochastic control of continuous processes. The novelty of this work is to incorporate intermediate expectation constraints on the canonical…

最优化与控制 · 数学 2020-04-22 Yuk-Loong Chow , Xiang Yu , Chao Zhou

We investigate the problem of best policy identification in discounted linear Markov Decision Processes in the fixed confidence setting under a generative model. We first derive an instance-specific lower bound on the expected number of…

机器学习 · 计算机科学 2022-08-12 Jerome Taupin , Yassir Jedra , Alexandre Proutiere

A dynamic mean field theory is developed for finite state and action Bayesian reinforcement learning in the large state space limit. In an analogy with statistical physics, the Bellman equation is studied as a disordered dynamical system;…

机器学习 · 统计学 2023-07-13 George Stamatescu

We study the convergence of Markov Decision Processes made of a large number of objects to optimization problems on ordinary differential equations (ODE). We show that the optimal reward of such a Markov Decision Process, satisfying a…

人工智能 · 计算机科学 2011-05-20 Nicolas Gast , Bruno Gaujal , Jean-Yves Le Boudec

We consider killed Markov decision processes for countable models on a finite time-interval. Existence of a uniform $\varepsilon$-optimal policy is proven. We show the correctness of the fundamental equation. The optimal control problem is…

最优化与控制 · 数学 2013-04-10 Nestor Parolya , Yaroslav Yeleyko

This paper investigates continuity properties of value functions and solutions for parametric optimization problems. These problems are important in operations research, control, and economics because optimality equations are their…

最优化与控制 · 数学 2021-09-15 Eugene A. Feinberg , Pavlo O. Kasyanov , David N. Kraemer

Robust Markov decision processes (MDPs) address the challenge of model uncertainty by optimizing the worst-case performance over an uncertainty set of MDPs. In this paper, we focus on the robust average-reward MDPs under the model-free…

机器学习 · 计算机科学 2023-05-19 Yue Wang , Alvaro Velasquez , George Atia , Ashley Prater-Bennette , Shaofeng Zou

Some approaches to solving challenging dynamic programming problems, such as Q-learning, begin by transforming the Bellman equation into an alternative functional equation, in order to open up a new line of attack. Our paper studies this…

最优化与控制 · 数学 2019-12-05 Qingyin Ma , John Stachurski

In this paper we propose a convex programming based method for computing robust regions of attraction for state-constrained perturbed discrete-time polynomial systems. The robust region of attraction of interest is a set of states such that…

动力系统 · 数学 2020-05-11 Bai Xue , Naijun Zhan , Yangjia Li

In this paper, we study the problem of optimizing the stability of positive semi-Markov jump linear systems. We specifically consider the problem of tuning the coefficients of the system matrices for maximizing the exponential decay rate of…

系统与控制 · 计算机科学 2020-09-22 Chengyan Zhao , Masaki Ogura , Kenji Sugimoto

Our first result is a statement of a somewhat general form of a non-substitution theorem for linear programming problems, along with a very easy proof of the same. Subsequently, we provide an easy proof of theorem 1 in a 1979 paper of Olvi…

最优化与控制 · 数学 2025-04-08 Somdeb Lahiri
‹ 上一页 1 8 9 10 下一页 ›