Related papers: Backward Stochastic Control System with Entropy Re…
In two-player zero-sum stochastic games, where two competing players make decisions under uncertainty, a pair of optimal strategies is traditionally described by Nash equilibrium and computed under the assumption that the players have…
We consider the problem of learning the optimal policy for Markov decision processes with safety constraints. We formulate the problem in a reach-avoid setup. Our goal is to design online reinforcement learning algorithms that ensure safety…
We examine a multi-stage stochastic optimization problem characterized by stagewise-independent, decision-dependent noises with strict constraints. The problem assumes convexity in that, following a specific relaxation, it transforms into a…
Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additional regularization bias. This work quantifies the impact of…
In this paper, we focus on a method based on optimal control to address the optimization problem. The objective is to find the optimal solution that minimizes the objective function. We transform the optimization problem into optimal…
This paper investigates the so-called reward-balancing methods, a novel class of algorithms for solving discounted-return reinforcement learning (RL) problems. These methods consist of iteratively adjusting the reward function to transform…
In this paper, we consider a class of stochastic control problems for stochastic differential equations with random coefficients. The control domain need not to be convex but the control process is not allowed to enter in diffusion term.…
This paper introduces a new formulation for stochastic optimal control and stochastic dynamic optimization that ensures safety with respect to state and control constraints. The proposed methodology brings together concepts such as…
In this paper, we consider optimal control of stochastic differential equations subject to an expected path constraint. The stochastic maximum principle is given for a general optimal stochastic control in terms of constrained FBSDEs. In…
While techniques have been developed for chance constrained stochastic optimal control using sample disturbance data that provide a probabilistic confidence bound for chance constraint satisfaction, far less is known about how to use sample…
We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework…
This paper is concerned with the maximum principle of stochastic optimal control problems, where the coefficients of the state equation and the cost functional are uncertain, and the system is generally under Markovian regime switching.…
When deploying artificial agents in real-world environments where they interact with humans, it is crucial that their behavior is aligned with the values, social norms or other requirements of that environment. However, many environments…
In this paper, we study the maximum principle for stochastic optimal control problems of forward-backward stochastic difference systems (FBS{\Delta}Ss). Two types of FBS{\Delta}Ss are investigated. The first one is described by a partially…
This paper presents a novel synthesis method for designing an optimal and robust guidance law for a non-throttleable upper stage of a launch vehicle, using a convex approach. In the unperturbed scenario, a combination of lossless and…
We establish an algorithm to learn feedback maps from data for a class of robust model predictive control (MPC) problems. The algorithm accounts for the approximation errors due to the learning directly at the synthesis stage, ensuring…
This paper investigates optimal control problems for delayed systems governed by Infinitely Anticipated Backward Stochastic Differential Equations (IABSDEs). Unlike existing frameworks limited to bounded delays, we introduce a generalized…
Probabilistic control design is founded on the principle that a rational agent attempts to match modelled with an arbitrary desired closed-loop system trajectory density. The framework was originally proposed as a tractable alternative to…
In this paper, we establish a general stochastic maximum principle for optimal control for systems described by a continuous-time Markov regime-switching stochastic recursive utilities model. The control domain is postulated not to be…
Inverse reinforcement learning aims to infer the reward function that explains expert behavior observed through trajectories of state--action pairs. A long-standing difficulty in classical IRL is the non-uniqueness of the recovered reward:…