中文
相关论文

相关论文: Bounding Fixed Points of Set-Based Bellman Operato…

200 篇论文

Motivated by uncertain parameters encountered in Markov decision processes (MDPs), we study the effect of parameter uncertainty on Bellman operator-based methods. Specifically, we consider a family of MDPs where the cost parameters are from…

最优化与控制 · 数学 2020-03-03 Sarah H. Q. Li , Assalé Adjé , Pierre-Loïc Garoche , Behçet Açıkmeşe

This paper analyzes finite state Markov Decision Processes (MDPs) with uncertain parameters in compact sets and re-examines results from robust MDP via set-based fixed point theory. To this end, we generalize the Bellman and policy…

机器学习 · 计算机科学 2023-08-09 Sarah H. Q. Li , Assalé Adjé , Pierre-Loïc Garoche , Behçet Açıkmeşe

We study the dynamic programming approach to revenue management in the context of attended home delivery. We draw on results from dynamic programming theory for Markov decision problems to show that the underlying Bellman operator has a…

最优化与控制 · 数学 2019-10-28 Denis Lebedev , Paul Goulart , Kostas Margellos

A Budgeted Markov Decision Process (BMDP) is an extension of a Markov Decision Process to critical applications requiring safety constraints. It relies on a notion of risk implemented in the shape of a cost signal constrained to lie below…

We establish the existence of fixed points for set-valued maps defined on metric spaces and satisfying a pointwise or a local version of Banach's contraction property. As an application, we demonstrate the existence of Nash equilibrium in a…

最优化与控制 · 数学 2022-11-24 Ted Loch-Temzelides

Learning and optimal control under robust Markov decision processes (MDPs) have received increasing attention, yet most existing theory, algorithms, and applications focus on finite-horizon or discounted models. Long-run average-reward…

最优化与控制 · 数学 2025-12-12 Shengbo Wang , Nian Si

The paper is concerned with two-person games with saddle point. We investigate the limits of value functions for long-time-average payoff, discounted average payoff, and the payoff that follows a probability density. Most of our assumptions…

最优化与控制 · 数学 2015-01-29 Dmitry Khlopin

We study sequential decision-making when the agent's internal model class is misspecified. Within the infinite-horizon Berk-Nash framework, stable behavior arises as a fixed point: the agent acts optimally relative to a subjective model,…

计算机科学与博弈论 · 计算机科学 2026-03-17 Quanyan Zhu , Zhengye Han

Markov decision processes (MDPs) with rewards are a widespread and well-studied model for systems that make both probabilistic and nondeterministic choices. A fundamental result about MDPs is that their minimal and maximal expected rewards…

计算机科学中的逻辑 · 计算机科学 2024-11-26 Kevin Batz , Benjamin Lucien Kaminski , Christoph Matheja , Tobias Winkler

In many multi-player interactions, players incur strictly positive costs each time they execute actions e.g. 'menu costs' or transaction costs in financial systems. Since acting at each available opportunity would accumulate prohibitively…

多智能体系统 · 计算机科学 2024-08-02 David Mguni

This paper studies value iteration for infinite horizon contracting Markov decision processes under convexity assumptions and when the state space is uncountable. The original value iteration is replaced with a more tractable form and the…

最优化与控制 · 数学 2018-02-21 Jeremy Yee

Markov decision processes (MDPs) are a popular model for performance analysis and optimization of stochastic systems. The parameters of stochastic behavior of MDPs are estimates from empirical observations of a system; their values are not…

人工智能 · 计算机科学 2017-10-26 Dimitri Scheftelowitsch , Peter Buchholz , Vahid Hashemi , Holger Hermanns

This paper considers mean field games in a multi-agent Markov decision process (MDP) framework. Each player has a continuum state and binary action, and benefits from the improvement of the condition of the overall population. Based on an…

最优化与控制 · 数学 2021-01-05 Minyi Huang , Yan Ma

In this paper, we consider a large class of constrained non-cooperative stochastic Markov games with countable state spaces and discounted cost criteria. In one-player case, i.e., constrained discounted Markov decision models, it is…

最优化与控制 · 数学 2021-12-16 Anna Jaśkiewicz , Andrzej S. Nowak

In this paper we develop a general framework to analyze stochastic dynamic problems with unbounded utility functions and correlated and unbounded shocks. We obtain new results of the existence and uniqueness of solutions to the Bellman…

理论经济学 · 经济学 2019-07-18 Juan Pablo Rincón-Zapatero

The solutions to many sequential decision-making problems are characterized by dynamic programming and Bellman's principle of optimality. However, due to the inherent complexity of solving Bellman's equation exactly, there has been…

系统与控制 · 电气工程与系统科学 2026-03-24 Bowen Li , Edwin K. P. Chong , Ali Pezeshki

We study learning dynamics induced by strategic agents who repeatedly play a game with an unknown payoff-relevant parameter. In each step, an information system estimates a belief distribution of the parameter based on the players'…

系统与控制 · 电气工程与系统科学 2020-10-20 Manxi Wu , Saurabh Amin , Asuman Ozdaglar

This paper studies function approximation for finite horizon discrete time Markov decision processes under certain convexity assumptions. Uniform convergence of these approximations on compact sets is proved under several sampling schemes…

最优化与控制 · 数学 2018-02-21 Jeremy Yee

We study the fixed point problem for a system of multivariate operators that are coordinate-wise monotone (i.e., nondecreasing or nonincreasing in each of the variables, independently), in the setting of quasi-ordered sets. We show that…

一般拓扑 · 数学 2012-09-03 Mircea-Dan Rus

We study value-iteration (VI) algorithms for solving general (a.k.a. multichain) Markov decision processes (MDPs) under the average-reward criterion, a fundamental but theoretically challenging setting. Beyond the difficulties inherent to…

最优化与控制 · 数学 2026-04-23 Matthew Zurek , Yudong Chen
‹ 上一页 1 2 3 10 下一页 ›