中文
相关论文

相关论文: Reduction of Markov Chains using a Value-of-Inform…

200 篇论文

This article presents several results establishing connections be- tween Markov chains and dynamical systems, from the point of view of open systems in physics. We show how all Markov chains can be understood as the information on one…

概率论 · 数学 2010-10-18 Stéphane Attal

We propose a novel randomized linear programming algorithm for approximating the optimal policy of the discounted Markov decision problem. By leveraging the value-policy duality and binary-tree data structures, the algorithm adaptively…

最优化与控制 · 数学 2019-06-04 Mengdi Wang

In networking applications, one often wishes to obtain estimates about the number of objects at different parts of the network (e.g., the number of cars at an intersection of a road network or the number of packets expected to reach a node…

社会与信息网络 · 计算机科学 2020-06-22 Harshal A. Chaudhari , Michael Mathioudakis , Evimaria Terzi

We consider the problem of learning a policy for a Markov decision process consistent with data captured on the state-actions pairs followed by the policy. We assume that the policy belongs to a class of parameterized policies which are…

最优化与控制 · 数学 2017-01-24 Manjesh K. Hanawal , Hao Liu , Henghui Zhu , Ioannis Ch. Paschalidis

The goal of this paper is to analyze distributional Markov Decision Processes as a class of control problems in which the objective is to learn policies that steer the distribution of a cumulative reward toward a prescribed target law,…

最优化与控制 · 数学 2026-02-09 Nicole Bäuerle , Athanasios Vasileiadis

Nonparametric identification and maximum likelihood estimation for finite-state hidden Markov models are investigated. We obtain identification of the parameters as well as the order of the Markov chain if the transition probability…

统计理论 · 数学 2015-10-01 Grigory Alexandrovich , Hajo Holzmann , Anna Leister

We consider state-aggregation schemes for Markov chains from an information-theoretic perspective. Specifically, we consider aggregating the states of a Markov chain such that the mutual information of the aggregated states separated by T…

物理与社会 · 物理学 2021-08-23 Mauro Faccin , Michael T. Schaub , Jean-Charles Delvenne

Modern safety-critical systems are heterogeneous, complex, and highly dynamic. They require reliability evaluation methods that go beyond the classical static methods such as fault trees, event trees, or reliability block diagrams.…

计算机科学中的逻辑 · 计算机科学 2020-04-15 Clemens Dubslaff , Andrey Morozov , Christel Baier , Klaus Janschek

This paper investigates MDPs with intermittent state information. We consider a scenario where the controller perceives the state information of the process via an unreliable communication channel. The transmissions of state information…

人工智能 · 计算机科学 2025-02-17 Gongpu Chen , Soung-Chang Liew

In this paper, we study a Markov chain-based stochastic gradient algorithm in general Hilbert spaces, aiming at approximating the optimal solution of a quadratic loss function. We establish probabilistic upper bounds on its convergence. We…

机器学习 · 统计学 2025-12-16 Priyanka Roy , Susanne Saminger-Platz

A method of constructing Markov chains on finite state spaces is provided. The chain is specified by three constraints: stationarity, dependence and marginal distributions. The generalized Pythagorean theorem in information geometry plays a…

统计理论 · 数学 2024-07-26 Tomonari Sei

We introduce a class of models for multidimensional control problems which we call skip-free Markov decision processes on trees. We describe and analyse an algorithm applicable to Markov decision processes of this type that are skip-free in…

最优化与控制 · 数学 2013-11-11 E. J. Collins

In this paper we provide faster algorithms for approximately solving discounted Markov Decision Processes in multiple parameter regimes. Given a discounted Markov Decision Process (DMDP) with $|S|$ states, $|A|$ actions, discount factor…

数据结构与算法 · 计算机科学 2020-12-24 Aaron Sidford , Mengdi Wang , Xian Wu , Yinyu Ye

Recently, Sidford, Wang, Wu and Ye (2018) developed an algorithm combining variance reduction techniques with value iteration to solve discounted Markov decision processes. This algorithm has a sublinear complexity when the discount factor…

最优化与控制 · 数学 2019-09-16 Marianne Akian , Stéphane Gaubert , Zheng Qu , Omar Saadi

We study the optimal liquidation problem in a market model where the bid price follows a geometric pure jump process whose local characteristics are driven by an unobservable finite-state Markov chain and by the liquidation rate. This model…

数理金融 · 定量金融 2019-06-27 Katia Colaneri , Zehra Eksi , Rüdiger Frey , Michaela Szölgyenyi

This work presents a low-rank tensor model for multi-dimensional Markov chains. A common approach to simplify the dynamical behavior of a Markov chain is to impose low-rankness on the transition probability matrix. Inspired by the success…

系统与控制 · 电气工程与系统科学 2024-11-05 Madeline Navarro , Sergio Rozada , Antonio G. Marques , Santiago Segarra

The paper is devoted to studies of perturbed Markov chains commonly used for description of information networks. In such models, the matrix of transition probabilities for the corresponding Markov chain is usually regularised by adding a…

Optimal Markov Decision Process policies for problems with finite state and action space are identified through a partial ordering by comparing the value function across states. This is referred to as state-based optimality. This paper…

最优化与控制 · 数学 2021-12-02 Dylan Solms

We present a Markov-chain analysis of blockwise-stochastic algorithms for solving partially block-separable optimization problems. Our main contributions to the extensive literature on these methods are statements about the Markov operators…

最优化与控制 · 数学 2023-11-01 D. Russell Luke

Robust Markov Decision Processes (MDPs) are a powerful framework for modeling sequential decision-making problems with model uncertainty. This paper proposes the first first-order framework for solving robust MDPs. Our algorithm interleaves…

最优化与控制 · 数学 2021-01-18 Julien Grand-Clément , Christian Kroer