Related papers: A variational formula for risk-sensitive reward
Robots that are trained to perform a task in a fixed environment often fail when facing unexpected changes to the environment due to a lack of exploration. We propose a principled way to adapt the policy for better exploration in changing…
In this paper, we study the stochastic combinatorial multi-armed bandit problem under semi-bandit feedback. While much work has been done on algorithms that optimize the expected reward for linear as well as some general reward functions,…
We propose a variational method to solve all three estimation problems for nonlinear stochastic dynamical systems: prediction, filtering, and smoothing. Our new approach is based upon a proper choice of cost function, termed the {\it…
This paper studies a one-sector optimal growth model with i.i.d. productivity shocks that are allowed to be unbounded. The utility function is assumed to be non-negative and unbounded from above. The novel feature in our framework is that…
A Markov decision problem is called reversible if the stationary controlled Markov chain is reversible under every stationary Markovian strategy. A natural application in which such problems arise is in the control of Metropolis-Hastings…
In this paper we investigate a kind of optimal control problem of coupled forward-backward stochastic system with jumps whose cost functional is defined through a coupled forward-backward stochastic differential equation with Brownian…
In this paper, we study a discrete-time stochastic optimal control problem under distribution uncertainty with convex control domain. By weak convergence method and Sion's minimax theorem, we obtain the variational inequality for cost…
We develop a method for computing policies in Markov decision processes with risk-sensitive measures subject to temporal logic constraints. Specifically, we use a particular risk-sensitive measure from cumulative prospect theory, which has…
We present a new, tractable method for solving and analyzing risk-aware control problems over finite and infinite, discounted time-horizons where the dynamics of the controlled process are described as a martingale problem. Supposing…
We develop an approach for solving time-consistent risk-sensitive stochastic optimization problems using model-free reinforcement learning (RL). Specifically, we assume agents assess the risk of a sequence of random variables using dynamic…
We develop a probabilistic framework for analysing model-based reinforcement learning in the episodic setting. We then apply it to study finite-time horizon stochastic control problems with linear dynamics but unknown coefficients and…
We study optimal control of Markov processes with age-dependent transition rates. The control policy is chosen continuously over time based on the state of the process and its age. We study infinite horizon discounted cost and infinite…
The importance of feedback control is being increasingly appreciated in quantum physics and applications. This paper describes the use of optimal control methods in the design of quantum feedback control systems, and in particular the paper…
A new approach to computation of optimal policies for MDP (Markov decision process) models is introduced. The main idea is to solve not one, but an entire family of MDPs, parameterized by a weighting factor $\zeta$ that appears in the…
In a Markovian stochastic volatility model, we consider financial agents whose investment criteria are modelled by forward exponential performance processes. The problem of contingent claim indifference valuation is first addressed and a…
We consider a parabolic optimal control problem with an initial measure control. The cost functional consists of a tracking term corresponding to the observation of the state at final time. Instead of a regularization term in the cost…
We consider the Merton problem of optimizing expected power utility of terminal wealth in the case of an unobservable Markov-modulated drift. What makes the model special is that the agent is allowed to purchase costly expert opinions of…
Model Predictive Control has emerged as a popular tool for robots to generate complex motions. However, the real-time requirement has limited the use of hard constraints and large preview horizons, which are necessary to ensure safety and…
Option-critic learning is a general-purpose reinforcement learning (RL) framework that aims to address the issue of long term credit assignment by leveraging temporal abstractions. However, when dealing with extended timescales, discounting…
Markov decision processes (MDPs) are widely used in modeling decision making problems in stochastic environments. However, precise specification of the reward functions in MDPs is often very difficult. Recent approaches have focused on…