English
Related papers

Related papers: Persistent-Transient Policy Evaluation for Markov …

200 papers

A Markov network characterizes the conditional independence structure, or Markov property, among a set of random variables. Existing work focuses on specific families of distributions (e.g., exponential families) and/or certain structures…

Machine Learning · Computer Science 2023-05-22 Yujia Zheng , Ignavier Ng , Yewen Fan , Kun Zhang

We build on a recently introduced geometric interpretation of Markov Decision Processes (MDPs) to analyze classical MDP-solving algorithms: Value Iteration (VI) and Policy Iteration (PI). First, we develop a geometry-based analytical…

Machine Learning · Computer Science 2025-03-07 Arsenii Mustafin , Aleksei Pakharev , Alex Olshevsky , Ioannis Ch. Paschalidis

The key assumption underlying linear Markov Decision Processes (MDPs) is that the learner has access to a known feature map $\phi(x, a)$ that maps state-action pairs to $d$-dimensional vectors, and that the rewards and transitions are…

Machine Learning · Computer Science 2023-09-20 Noah Golowich , Ankur Moitra , Dhruv Rohatgi

This paper presents different recursive formulas for computing the marginals and the normalizing constant of a Gibbs distribution $\pi$: The common thread is the use of the underlying Markov properties of such processes. The procedures are…

Probability · Mathematics 2025-11-07 Cécile Hardouin , Xavier Guyon

We consider the control of a Markov decision process (MDP) that undergoes an abrupt change in its transition kernel (mode). We formulate the problem of minimizing regret under control-switching based on mode change detection, compared to a…

Systems and Control · Electrical Eng. & Systems 2022-10-11 Nathan Dahlin , Subhonmesh Bose , Venugopal V. Veeravalli

We present a flexible Bayesian semiparametric mixed model for longitudinal data analysis in the presence of potentially high-dimensional categorical covariates. Building on a novel hidden Markov tensor decomposition technique, our proposed…

Methodology · Statistics 2022-08-05 Giorgio Paulon , Peter Müller , Abhra Sarkar

We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove…

Machine Learning · Computer Science 2011-05-02 Shie Mannor , John Tsitsiklis

Motivated from Bertsekas' recent study on policy iteration (PI) for solving the problems of infinite-horizon discounted Markov decision processes (MDPs) in an on-line setting, we develop an off-line PI integrated with a multi-policy…

Optimization and Control · Mathematics 2021-12-07 Hyeong Soo Chang

Principal component analysis (PCA) is a well-established method commonly used to explore and visualise data. A classical PCA model is the fixed effect model where data are generated as a fixed structure of low rank corrupted by noise. Under…

Methodology · Statistics 2013-05-13 Marie Verbanck , Julie Josse , François Husson

We study quasi-stationary distributions and quasi-limiting behavior of Markov chains in general reducible state spaces with absorption. We propose a set of assumptions dealing with particular situations where the state space can be…

Probability · Mathematics 2026-01-14 Nicolas Champagnat , Denis Villemonais

We consider a model for linear transient price impact for multiple assets that takes cross-asset impact into account. Our main goal is to single out properties that need to be imposed on the decay kernel so that the model admits…

Trading and Market Microstructure · Quantitative Finance 2015-09-10 Aurélien Alfonsi , Alexander Schied , Florian Klöck

We present a learning model predictive control (MPC) scheme for chance-constrained Markov jump systems with unknown switching probabilities. Using samples of the underlying Markov chain, ambiguity sets of transition probabilities are…

Optimization and Control · Mathematics 2023-01-06 Mathijs Schuurmans , Panagiotis Patrinos

Transitional inference is an empiricism concept, rooted and practiced in clinical medicine since ancient Greece. Knowledge and experiences gained from treating one entity are applied to treat a related but distinctively different one. This…

Statistics Theory · Mathematics 2020-12-15 Xinran Li , Xiao-Li Meng

In this paper, we consider general Markov chains (MC), specified by the transition probability (kernel) $ P (x, E) $, finitely additive in the second argument. Such MC are studied within the framework of the functional operator treatment.…

Probability · Mathematics 2022-01-11 Alexander Zhdanok , Anna Khuruma

Markov decision processes (MDPs) with rewards are a widespread and well-studied model for systems that make both probabilistic and nondeterministic choices. A fundamental result about MDPs is that their minimal and maximal expected rewards…

Logic in Computer Science · Computer Science 2024-11-26 Kevin Batz , Benjamin Lucien Kaminski , Christoph Matheja , Tobias Winkler

The problem of covariance estimation for replicated surface-valued processes is examined from the functional data analysis perspective. Considerations of statistical and computational efficiency often compel the use of separability of the…

Methodology · Statistics 2021-10-25 Tomas Masak , Victor M. Panaretos

While there is an extensive body of research analyzing policy gradient methods for discounted cumulative-reward MDPs, prior work on policy gradient methods for average-reward MDPs has been limited, with most existing results restricted to…

Optimization and Control · Mathematics 2026-02-23 Jongmin Lee , Ernest K. Ryu

For a discrete-time Markov chain $\{X(t)\}$ evolving on $\Re^\ell$ with transition kernel $P$, natural, general conditions are developed under which the following are established: 1. The transition kernel $P$ has a purely discrete spectrum,…

Probability · Mathematics 2019-07-19 Adithya Devraj , Ioannis Kontoyiannis , Sean Meyn

We study positive recurrence and transience of a two-station network in which the behavior of the server in each station is governed by a Markov chain with a finite number of server states; this service process can represent various service…

Probability · Mathematics 2013-08-29 Toshihisa Ozawa

Markov Decision Processes (MDPs) have been used to formulate many decision-making problems in science and engineering. The objective is to synthesize the best decision (action selection) policies to maximize expected rewards (or minimize…

Optimization and Control · Mathematics 2015-07-07 Mahmoud El Chamie , Behcet Acikmese