中文
相关论文

相关论文: Unbounded Markov Dynamic Programming with Weighted…

200 篇论文

This paper shows the usefulness of Perov's contraction principle, which generalizes Banach's contraction principle to a vector-valued metric, for studying dynamic programming problems in which the discount factor can be stochastic. The…

理论经济学 · 经济学 2021-09-10 Alexis Akira Toda

This paper deals with unconstrained discounted continuous-time Markov decision processes in Borel state and action spaces. Under some conditions imposed on the primitives, allowing unbounded transition rates and unbounded (from both above…

最优化与控制 · 数学 2011-03-02 Alexey Piunovskiy , Yi Zhang

The problem of solving Markov decision processes under function approximation remains a fundamental challenge, even under linear function approximation settings. A key difficulty arises from a geometric mismatch: while the Bellman…

机器学习 · 计算机科学 2026-04-09 Hyukjun Yang , Han-Dong Lim , Donghwan Lee

We consider a discrete-time Markov decision process with Borel state and action spaces. The performance criterion is to maximize a total expected {utility determined by unbounded return function. It is shown the existence of optimal…

概率论 · 数学 2018-10-08 François Dufour , Alexandre Genadot

We propose a new approach to solving dynamic decision problems with rewards that are unbounded below. The approach involves transforming the Bellman equation in order to convert an unbounded problem into a bounded one. The major advantage…

理论经济学 · 经济学 2019-12-02 Qingyin Ma , John Stachurski

The paper deals with a risk averse dynamic programming problem with infinite horizon. First, the required assumptions are formulated to have the problem well defined. Then the Bellman equation is derived, which may be also seen as a…

最优化与控制 · 数学 2022-08-04 Martin Šmíd , Miloš Kopa

Learning and optimal control under robust Markov decision processes (MDPs) have received increasing attention, yet most existing theory, algorithms, and applications focus on finite-horizon or discounted models. Long-run average-reward…

最优化与控制 · 数学 2025-12-12 Shengbo Wang , Nian Si

This paper investigates discrete-time Markov decision processes with recursive utilities (or payoffs) defined by the classic CES aggregator and the Kreps-Porteus certainty equivalent operator. According to the classification introduced by…

最优化与控制 · 数学 2025-07-11 Anna Jaśkiewicz , Andrzej S. Nowak

Finding optimal policies which maximize long term rewards of Markov Decision Processes requires the use of dynamic programming and backward induction to solve the Bellman optimality equation. However, many real-world problems require…

机器学习 · 计算机科学 2023-01-10 Mridul Agarwal , Vaneet Aggarwal

In this paper we develop a general framework to analyze stochastic dynamic problems with unbounded utility functions and correlated and unbounded shocks. We obtain new results of the existence and uniqueness of solutions to the Bellman…

理论经济学 · 经济学 2019-07-18 Juan Pablo Rincón-Zapatero

We study the dynamic programming approach to revenue management in the context of attended home delivery. We draw on results from dynamic programming theory for Markov decision problems to show that the underlying Bellman operator has a…

最优化与控制 · 数学 2019-10-28 Denis Lebedev , Paul Goulart , Kostas Margellos

The problem of constrained Markov decision process is considered. An agent aims to maximize the expected accumulated discounted reward subject to multiple constraints on its costs (the number of constraints is relatively small). A new dual…

In this paper, we study a Markov decision process with a non-linear discount function and with a Borel state space. We define a recursive discounted utility, which resembles non-additive utility functions considered in a number of models in…

最优化与控制 · 数学 2025-10-16 Nicole Bäuerle , Anna Jaśkiewicz , Andrzej S. Nowak

We study optimality for the safety-constrained Markov decision process which is the underlying framework for safe reinforcement learning. Specifically, we consider a constrained Markov decision process (with finite states and finite…

系统与控制 · 电气工程与系统科学 2023-07-13 Rahul Misra , Rafał Wisniewski , Carsten Skovmose Kallesøe

We consider the optimal stopping problem consisting in, given a strong Markov process, a reward function and a discount rate, finding the stopping time such that the expected reward at the stopping time is maximum. The approach we follow,…

概率论 · 数学 2014-05-30 Fabián Crocce

In this work, we study discrete-time Markov decision processes (MDPs) under constraints with Borel state and action spaces and where all the performance functions have the same form of the expected total reward (ETR) criterion over the…

概率论 · 数学 2019-05-10 F. Dufour , Alexandre Genadot

A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…

机器学习 · 计算机科学 2023-09-04 Falcon Z. Dai

We study a Markov decision problem in which the state space is the set of finite marked point configurations in the plane, the actions represent thinnings, the reward is proportional to the mark sum which is discounted over time, and the…

概率论 · 数学 2023-09-08 M. N. M. van Lieshout

We consider non-standard Markov Decision Processes (MDPs) where the target function is not only a simple expectation of the accumulated reward. Instead, we consider rather general functionals of the joint distribution of terminal state and…

最优化与控制 · 数学 2025-10-16 Nicole Bäuerle , Tamara Göll , Anna Jaśkiewicz

Markov decision processes (MDPs) are used to model a wide variety of applications ranging from game playing over robotics to finance. Their optimal policy typically maximizes the expected sum of rewards given at each step of the decision…

机器学习 · 计算机科学 2025-05-26 Maximilian Nägele , Jan Olle , Thomas Fösel , Remmy Zen , Florian Marquardt
‹ 上一页 1 2 3 10 下一页 ›