Related papers: Long-Run Average Reward Maximization of A Regulate…
This paper describes a novel method to solve average-reward semi-Markov decision processes, by reducing them to a minimal sequence of cumulative reward problems. The usual solution methods for this type of problems update the gain (optimal…
This paper investigates the so-called reward-balancing methods, a novel class of algorithms for solving discounted-return reinforcement learning (RL) problems. These methods consist of iteratively adjusting the reward function to transform…
In this paper, we study optimal control problems for multiclass GI/M/n+M queues in an alternating renewal (up-down) random environment in the Halfin-Whitt regime. Assuming that the downtimes are asymptotically negligible and only the…
We propose a reinforcement learning (RL) framework for multi-objective decision-making, where the agent seeks to optimize a vector of rewards rather than a single scalar value. The objective is to ensure that the time-averaged reward vector…
Revisiting the continuous-time Mean-Variance (MV) Portfolio Optimization problem, we model the market dynamics with a jump-diffusion process and apply Reinforcement Learning (RL) techniques to facilitate informed exploration within the…
This paper presents an efficient procedure for multi-objective model checking of long-run average reward (aka: mean pay-off) and total reward objectives as well as their combination. We consider this for Markov automata, a compositional…
This paper is concerned with the maximum principle of stochastic optimal control problems, where the coefficients of the state equation and the cost functional are uncertain, and the system is generally under Markovian regime switching.…
We develop a neural-network framework for multi-period risk--reward stochastic control problems with constrained two-step feedback policies that may be discontinuous in the state. We allow a broad class of objectives built on a…
We develop a comprehensive mathematical and computational framework for optimal production planning in economies governed by stochastic regime switches driven by a continuous-time Markov chain. The value functions of the underlying…
We consider a diffusion risk model where dividends are paid at rate $U(t) \in [0, u_0]$. We are interested in maximising the dividend payments under a drawdown constraint, that is, we penalise a drawdown size larger than a level $d > 0$. We…
We consider a discrete-time dividend payout problem with risk sensitive shareholders. It is assumed that they are equipped with a risk aversion coefficient and construct their discounted payoff with the help of the exponential premium…
When randomness in demand affects the sales of a product, retailers use dynamic pricing strategies to maximize their profits. In this article, we formulate the pricing problem as a continuous-time stochastic optimal control problem and find…
Dynamic and evolving operational and economic environments present significant challenges for decision-making. We explore a simulation optimization problem characterized by non-stationary input distributions with regime-switching dynamics…
This work concerns controlled Markov chains with finite state and action spaces. The transition law satisfies the simultaneous Doeblin condition, and the performance of a control policy is measured by the (long-run) risk-sensitive average…
By adopting a distributional viewpoint on law-invariant convex risk measures, we construct dynamics risk measures (DRMs) at the distributional level. We then apply these DRMs to investigate Markov decision processes, incorporating latent…
Stochastic domains often involve risk-averse decision makers. While recent work has focused on how to model risk in Markov decision processes using risk measures, it has not addressed the problem of solving large risk-averse formulations.…
We consider a jump-diffusion mean field control problem with regime switching in the state dynamics. The corresponding value function is characterized as the unique viscosity solution of a HJB master equation on the space of probability…
This paper considers the portfolio management problem of optimal investment, consumption and life insurance. We are concerned with time inconsistency of optimal strategies. Natural assumptions, like different discount rates for consumption…
We consider the optimal dividend problem under a habit formation constraint that prevents the dividend rate to fall below a certain proportion of its historical maximum, the so-called drawdown constraint. This is an extension of the optimal…
We consider a class of diffusions controlled through the drift and jump size, and driven by a jump L\'evy process and a nondegenerate Wiener process, and we study infinite horizon (ergodic) risk-sensitive control problem for this model. We…