Related papers: The Infinite-Dimensional Standard and Strict Bound…
We provide sufficient conditions for the existence of invariant probability measures for generic stochastic differential equations with finite time delay. This is achieved by means of the Krylov-Bogoliubov method. Furthermore, we focus on…
Safe exploration remains a fundamental challenge in reinforcement learning (RL), limiting the deployment of RL agents in the real world. We propose Sampling-Based Safe Reinforcement Learning (SBSRL), a model-based RL algorithm that…
This article is a survey of results concerning an inequality, which may be seen as a versatile tool to solve problems in the domain of Applied Probability. The inequality, which we call BRS-inequality, gives a convenient upper bound for the…
Batch reinforcement learning (RL) aims at leveraging pre-collected data to find an optimal policy that maximizes the expected total rewards in a dynamic environment. The existing methods require absolutely continuous assumption (e.g., there…
Estimation of solution norms and stability for time-dependent nonlinear systems is ubiquitous in numerous engineering, natural science and control problems. Yet, practically valuable results are rare in this area. This paper develops a…
Designing model-free algorithms for distributionally robust reinforcement learning (DRRL) poses fundamental challenges. The robust Bellman operator is nonlinear in the transition kernel, which makes one-sample Bellman updates biased, while…
Hamilton-Jacobi reachability (HJR) provides a value function that encodes the set of states from which a system with bounded control inputs can reach or avoid a target despite any bounded disturbance, and the corresponding robust, optimal…
Representation learning (RL) methods learn objects' latent embeddings where information is preserved by distances. Since distances are invariant to certain linear transformations, one may obtain different embeddings while preserving the…
In this work, we propose a methodology for the expression of necessary and sufficient Lyapunov-like conditions for the existence of stabilizing feedback laws. The methodology is an extension of the well-known Control Lyapunov Function (CLF)…
Deep reinforcement learning has been recognized as a promising tool to address the challenges in real-time control of power systems. However, its deployment in real-world power systems has been hindered by a lack of explicit stability and…
This paper develops a model-based reinforcement learning (MBRL) framework for learning online the value function of an infinite-horizon optimal control problem while obeying safety constraints expressed as control barrier functions (CBFs).…
In this paper we apply the formalism of translation invariant (continuous) matrix product states in the thermodynamic limit to $(1+1)$ dimensional critical models. Finite bond dimension bounds the entanglement entropy and introduces an…
Two main challenges in Reinforcement Learning (RL) are designing appropriate reward functions and ensuring the safety of the learned policy. To address these challenges, we present a theoretical framework for Inverse Reinforcement Learning…
We study integral-to-integral input-to-state stability for infinite-dimensional linear systems with inputs and trajectories in $L^p$-spaces. We start by developing the corresponding admissibility theory for linear systems with unbounded…
In this article we are interested in the boundary stabilization in finite time of one-dimensional linear hyperbolic balance laws with coefficients depending on time and space. We extend the so called "backstepping method" by introducing…
A differential dynamic programming (DDP)-based framework for inverse reinforcement learning (IRL) is introduced to recover the parameters in the cost function, system dynamics, and constraints from demonstrations. Different from existing…
The field of risk-constrained reinforcement learning (RCRL) has been developed to effectively reduce the likelihood of worst-case scenarios by explicitly handling risk-measure-based constraints. However, the nonlinearity of risk measures…
Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to settings such as verifiable rewards and multi-objective alignment, RMs are expected to encode…
We introduce a novel framework for analyzing reinforcement learning (RL) in continuous state-action spaces, and use it to prove fast rates of convergence in both off-line and on-line settings. Our analysis highlights two key stability…
We study inventory control policies for pharmaceutical supply chains, addressing challenges such as perishability, yield uncertainty, and non-stationary demand, combined with batching constraints, lead times, and lost sales. Collaborating…