Related papers: Long-Run Average Reward Maximization of A Regulate…
In this paper we study a class of optimal dividend and investment problems assuming that the underlying reserve process follows the Sparre Andersen model, that is, the claim frequency is a "renewal" process, rather than a standard compound…
This paper studies a type of rank-based mean field game in which competing agents strategically switch among multiple effort regimes. We propose an entropy regularized auxiliary problem where the switching decisions are randomized to the…
We develop an approach for two player constraint zero-sum and nonzero-sum stochastic differential games, which are modeled by Markov regime-switching jump-diffusion processes. We provide the relations between a usual stochastic optimal…
We present a case study applying learning-based distributionally robust model predictive control to highway motion planning under stochastic uncertainty of the lane change behavior of surrounding road users. The dynamics of road users are…
This contribution mainly focuses on the finite horizon optimal control problems of a susceptible-infected-vaccinated(SIV) epidemic system governed by reaction-diffusion equations and Markov switching. Stochastic dynamic programming is…
In this work, we consider the optimal portfolio selection problem under hard constraints on trading volume amounts when the dynamics of the risky asset returns are governed by a discrete-time approximation of the Markov-modulated geometric…
Stochastic optimal control problems have a long tradition in applied probability, with the questions addressed being of high relevance in a multitude of fields. Even though theoretical solutions are well understood in many scenarios, their…
Option-critic learning is a general-purpose reinforcement learning (RL) framework that aims to address the issue of long term credit assignment by leveraging temporal abstractions. However, when dealing with extended timescales, discounting…
This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type of exploratory formulation under entropy regularization where the agent randomizes both the timing…
We consider a portfolio optimization problem in a defaultable market with finitely-many economical regimes, where the investor can dynamically allocate her wealth among a defaultable bond, a stock, and a money market account. The market…
We consider the problem of controlling a Markov decision process (MDP) with a large state space, so as to minimize average cost. Since it is intractable to compete with the optimal policy for large scale problems, we pursue the more modest…
This paper deals with discrete-time Markov control processes on a general state space. A long-run risk-sensitive average cost criterion is used as a performance measure. The one-step cost function is nonnegative and possibly unbounded.…
The optimization criterion for dividends from a risky business is most often formalized in terms of the expected present value of future dividends. That criterion disregards a potential, explicit demand for stability of dividends. In…
This paper deals with numerical solutions of maximizing expected utility from terminal wealth under a non-bankruptcy constraint. The wealth process is subject to shocks produced by a general marked point process. The problem of the agent is…
We consider a singular control problem with regime switching that arises in problems of optimal investment decisions of cash-constrained firms. The value function is proved to be the unique viscosity solution of the associated…
We introduce the Lyapunov approach to optimal control problems of average risk-sensitive Markov control processes with general risk maps. Motivated by applications in particular to behavioral economics, we consider possibly non-convex risk…
This paper addresses objectives tailored to the risk-averse optimization of accumulated rewards in Markov decision processes (MDPs). The studied objectives require maximizing the expected value of the accumulated rewards minus a penalty…
This paper analyzes and explicitly solves a class of long-term average impulse control problems and a related class of singular control problems. The underlying process is a general one-dimensional diffusion with appropriate boundary…
This paper considers an optimal control of a big financial company with debt liability under bankrupt probability constraints. The company, which faces constant liability payments and has choices to choose various production/business…
We study a an optimal high frequency trading problem within a market microstructure model designed to be a good compromise between accuracy and tractability. The stock price is driven by a Markov Renewal Process (MRP), while market orders…