Related papers: On Bellman's principle with inequality constraints
We study optimality for the safety-constrained Markov decision process which is the underlying framework for safe reinforcement learning. Specifically, we consider a constrained Markov decision process (with finite states and finite…
While the standard approach to quantum systems studies length preserving linear transformations of wave functions, the Markov picture focuses on trace preserving operators on the space of Hermitian (self-adjoint) matrices. The Markov…
We study a limit behavior of a sequence of Markov processes (or Markov chains) such that their distributions outside of any neighborhood of a "singular" point attract to some probability law. In any neighborhood of this point the behavior…
The idea of the restricted mean has been used to establish a significantly improved version of Markov's inequality that does not require any new assumptions. The result immediately extends on Chebyshev's inequalities and Chernoff's bound.…
Both constrained and unconstrained optimization problems regularly appear in recursive tracking problems engineers currently address -- however, constraints are rarely exploited for these applications. We define the Kalman Filter and…
We consider a Markov decision process subject to model uncertainty in a Bayesian framework, where we assume that the state process is observed but its law is unknown to the observer. In addition, while the state process and the controls are…
We show that the ability to consider counterfactual situations is a necessary assumption of Bell's theorem, and that, to allow Bell inequality violations while maintaining all other assumptions, we just require certain measurement choices…
The problem of constrained Markov decision process is considered. An agent aims to maximize the expected accumulated discounted reward subject to multiple constraints on its costs (the number of constraints is relatively small). A new dual…
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the gradient of the objective.
We investigate discrete-time mean-variance portfolio selection problems viewed as a Markov decision process. We transform the problems into a new model with deterministic transition function for which the Bellman optimality equation holds.…
The study of Markov models is central to control theory and machine learning. A quantum analogue of partially observable Markov decision process was studied in (Barry, Barry, and Aaronson, Phys. Rev. A, 90, 2014). It was proved that…
We explore the challenges posed by the violation of Bell-like inequalities by $d$-dimensional systems exposed to imperfect state-preparation and measurement settings. We address, in particular, the limit of high-dimensional systems,…
This paper studies value iteration for infinite horizon contracting Markov decision processes under convexity assumptions and when the state space is uncountable. The original value iteration is replaced with a more tractable form and the…
An abstract treatment of Bell inequalities is proposed, in which the parameters characterizing Bell's observable can be times rather than directions. The violation of a Bell inequality might then be taken to mean that a property of a system…
Under the expected total reward criterion, the optimal value of a finite-horizon Markov decision process can be determined by solving the Bellman equations. The equations were extended by D. J. White to processes with vector rewards in…
Proposals for Bell inequality tests on systems restricted by superselection rules often require operations that are difficult to implement in practice. In this paper, we derive a new Bell inequality, where pairs of states are used to…
The Bell inequalities can be violated by postselecting on the results of a measurement of the Bell states. If information about the original state preparation is available, we point out how the violation can be reproduced classically by…
In this work we address the problem of finding feasible policies for Constrained Markov Decision Processes under probability one constraints. We argue that stationary policies are not sufficient for solving this problem, and that a rich…
We formally prove the existence of an enduring incongruence pervading a widespread interpretation of the Bell inequality and explain how to rationally avoid it with a natural assumption justified by explicit reference to a mathematical…
We study Markov decision problems where the agent does not know the transition probability function mapping current states and actions to future states. The agent has a prior belief over a set of possible transition functions and updates…