中文
相关论文

相关论文: Online Markov decision processes with Kullback-Lei…

200 篇论文

This paper is devoted to solving a time-inconsistent risk-sensitive control problem with parameter $\e$ and its limit case ($\e\rightarrow0^+$) for countable-stated Markov decision processes (MDPs for short). Since the cost functional is…

最优化与控制 · 数学 2020-10-22 Hongwei Mei

This paper studies a finite-horizon Markov decision problem with information-theoretic constraints, where the goal is to minimize directed information from the controlled source process to the control process, subject to stage-wise cost…

系统与控制 · 电气工程与系统科学 2025-09-04 Zixuan He , Charalambos D. Charalambous , Photios A. Stavrou

The objective of this work is to study continuous-time Markov decision processes on a general Borel state space with both impulsive and continuous controls for the infinite-time horizon discounted cost. The continuous-time controlled…

最优化与控制 · 数学 2019-08-17 François Dufour , Alexei Piunovskiy

We investigate the problem of best-policy identification in discounted Markov Decision Processes (MDPs) when the learner has access to a generative model. The objective is to devise a learning algorithm returning the best policy as early as…

机器学习 · 统计学 2021-05-11 Aymen Al Marjani , Alexandre Proutiere

Recently, continual learning has received a lot of attention. One of the significant problems is the occurrence of \emph{concept drift}, which consists of changing probabilistic characteristics of the incoming data. In the case of the…

机器学习 · 计算机科学 2022-10-11 Sebastián Basterrech , Michal Woźniak

Explaining adaptive behavior is a central problem in artificial intelligence research. Here we formalize adaptive agents as mixture distributions over sequences of inputs and outputs (I/O). Each distribution of the mixture constitutes a…

人工智能 · 计算机科学 2009-12-31 Pedro A. Ortega , Daniel A. Braun

This paper studies distributed Q-learning for Linear Quadratic Regulator (LQR) in a multi-agent network. The existing results often assume that agents can observe the global system state, which may be infeasible in large-scale systems due…

多智能体系统 · 计算机科学 2020-12-24 Hang Wang , Sen Lin , Hamid Jafarkhani , Junshan Zhang

A new approach to computation of optimal policies for MDP (Markov decision process) models is introduced. The main idea is to solve not one, but an entire family of MDPs, parameterized by a weighting factor $\zeta$ that appears in the…

最优化与控制 · 数学 2018-09-18 Ana Bušić , Sean Meyn

This work concerns controlled Markov chains with finite state and action spaces. The transition law satisfies the simultaneous Doeblin condition, and the performance of a control policy is measured by the (long-run) risk-sensitive average…

概率论 · 数学 2007-05-23 Rolando Cavazos-Cadena , Daniel Hernandez-Hernandez

We consider the problem of online learning of optimal control for repeatedly operated systems in the presence of parametric uncertainty. During each round of operation, environment selects system parameters according to a fixed but unknown…

机器学习 · 计算机科学 2016-09-20 Theja Tulabandhula

Measuring states in reinforcement learning (RL) can be costly in real-world settings and may negatively influence future outcomes. We introduce the Actively Observable Markov Decision Process (AOMDP), where an agent not only selects control…

机器学习 · 计算机科学 2025-10-17 Daiqi Gao , Ziping Xu , Aseel Rawashdeh , Predrag Klasnja , Susan A. Murphy

Suppose an online platform wants to compare a treatment and control policy, e.g., two different matching algorithms in a ridesharing system, or two different inventory management algorithms in an online retail site. Standard randomized…

统计方法学 · 统计学 2022-12-27 Peter Glynn , Ramesh Johari , Mohammad Rasouli

This paper considers a multiagent, connected, robotic fleet where the primary functionality of the agents is sensing. A distributed multi-sensor control strategy maximizes the value of the collective sensing capability of the fleet, using…

多智能体系统 · 计算机科学 2022-03-04 Tianqi Li , Lucas W. Krakow , Swaminathan Gopalswamy

We study an online learning problem on dynamic pricing and resource allocation, where we make joint pricing and inventory decisions to maximize the overall net profit. We consider the stochastic dependence of demands on the price, which…

机器学习 · 计算机科学 2025-05-23 Jianyu Xu , Xuan Wang , Yu-Xiang Wang , Jiashuo Jiang

We consider a finite-state, continuous-time Markov process, represented in the "linear framework" by a directed graph with labelled edges which specifies the infinitesimal generator of the process. If the graph is strongly connected, the…

生物物理 · 物理学 2023-10-17 Ugur Cetiner , Jeremy Gunawardena

We consider the problem of learning a policy for a Markov decision process consistent with data captured on the state-actions pairs followed by the policy. We assume that the policy belongs to a class of parameterized policies which are…

最优化与控制 · 数学 2017-01-24 Manjesh K. Hanawal , Hao Liu , Henghui Zhu , Ioannis Ch. Paschalidis

This paper studies motion planning of a mobile robot under uncertainty. The control objective is to synthesize a {finite-memory} control policy, such that a high-level task specified as a Linear Temporal Logic (LTL) formula is satisfied…

机器人学 · 计算机科学 2017-10-24 Meng Guo , Michael M. Zavlanos

This note re-visits the rolling-horizon control approach to the problem of a Markov decision process (MDP) with infinite-horizon discounted expected reward criterion. Distinguished from the classical value-iteration approach, we develop an…

最优化与控制 · 数学 2022-06-07 Hyeong Soo Chang

Many processes, such as discrete event systems in engineering or population dynamics in biology, evolve in discrete space and continuous time. We consider the problem of optimal decision making in such discrete state and action space…

机器学习 · 计算机科学 2020-10-27 Bastian Alt , Matthias Schultheis , Heinz Koeppl

The paper considers a class of multi-agent Markov decision processes (MDPs), in which the network agents respond differently (as manifested by the instantaneous one-stage random costs) to a global controlled state and the control actions of…

机器学习 · 统计学 2015-06-04 Soummya Kar , Jose' M. F. Moura , H. Vincent Poor