Related papers: The Value Functions of Markov Decision Processes
We study Markov-modulated affine processes (abbreviated MMAPs), a class of Markov processes that are created from affine processes by allowing some of their coefficients to be a function of an exogenous Markov process. MMAPs allow for…
Two approaches to studying the correlation functions of the binary Markov sequences are considered. The first of them is based on the study of probability of occurring different ''words'' in the sequence. The other one uses recurrence…
A study of time homogeneous, real valued Markov processes with a special property and a non-atomic initial distribution is provided. The new notion of a function of evolution of distribution which determines the dependency between one…
Markov decision processes (MDP) are a well-established model for sequential decision-making in the presence of probabilities. In robust MDP (RMDP), every action is associated with an uncertainty set of probability distributions, modelling…
Inference, prediction and control of complex dynamical systems from time series is important in many areas, including financial markets, power grid management, climate and weather modeling, or molecular dynamics. The analysis of such highly…
The value 1 problem is a decision problem for probabilistic automata over finite words: are there words accepted by the automaton with arbitrarily high probability? Although undecidable, this problem attracted a lot of attention over the…
This paper presents an axiomatic approach to finite Markov decision processes where the discount rate is zero. One of the principal difficulties in the no discounting case is that, even if attention is restricted to stationary policies, a…
We incorporate safety specifications into dynamic programming. Explicitly, we address the minimization problem of a Markov decision process up to a stopping time with safety constraints. To incorporate safety into dynamic programming, we…
We develop a method for computing policies in Markov decision processes with risk-sensitive measures subject to temporal logic constraints. Specifically, we use a particular risk-sensitive measure from cumulative prospect theory, which has…
The paper deals with some elementary problems about various mean value properties and their connections to harmonic functions and random walks.
Empirical processes for stationary, causal sequences are considered. We establish empirical central limit theorems for classes of indicators of left half lines, absolutely continuous functions and piecewise differentiable functions. Sample…
Aggregation functions are generally defined and used to combine several numerical values into a single one, so that the final result of the aggregation takes into account all the individual values in a given manner. Such functions are…
We introduce a general method for the study of memory in symbolic sequences based on higher-order Markov analysis. The Markov process that best represents a sequence is expressed as a mixture of matrices of minimal orders, enabling the…
In many real-world applications, the reward function is too complex to be manually specified. In such cases, reward functions must instead be learned from human feedback. Since the learned reward may fail to represent user preferences, it…
Our goal is to develop a partial ordering method for comparing stochastic choice functions on the basis of their individual rationality. To this end, we assign to any stochastic choice function a one-parameter class of deterministic choice…
Markov decision models (MDM) used in practical applications are most often less complex than the underlying `true' MDM. The reduction of model complexity is performed for several reasons. However, it is obviously of interest to know what…
We study the problem of merging sequential or independent e-values into one e-value or e-process. We describe a class of e-value merging functions via martingales and show that it dominates all merging methods for sequential e-values. All…
We consider statistical Markov Decision Processes where the decision maker is risk averse against model ambiguity. The latter is given by an unknown parameter which influences the transition law and the cost functions. Risk aversion is…
Weighted mean value identities over balls are considered for harmonic functions and their derivatives. Logarithmic and other weights are involved in these identities for functions. Some applications of weighted identities are presented.…
We obtain weak rates for approximation of an integral functional of a Markov process by integral sums. An assumption on the process is formulated only in terms of its transition probability density, and, therefore, our approach is not…