中文
相关论文

相关论文: The Infinite-Dimensional Standard and Strict Bound…

200 篇论文

We provide sufficient conditions for the existence of invariant probability measures for generic stochastic differential equations with finite time delay. This is achieved by means of the Krylov-Bogoliubov method. Furthermore, we focus on…

动力系统 · 数学 2026-05-15 Mark van den Bosch , Onno van Gaans , Sjoerd Verduyn Lunel

Safe exploration remains a fundamental challenge in reinforcement learning (RL), limiting the deployment of RL agents in the real world. We propose Sampling-Based Safe Reinforcement Learning (SBSRL), a model-based RL algorithm that…

机器学习 · 计算机科学 2026-05-20 Luca Vignola , Bruce D. Lee , Manish Prajapat , Manuel Wendl , Melanie Zeilinger , Andreas Krause , Yarden As

This article is a survey of results concerning an inequality, which may be seen as a versatile tool to solve problems in the domain of Applied Probability. The inequality, which we call BRS-inequality, gives a convenient upper bound for the…

概率论 · 数学 2020-07-13 F. Thomas Bruss

Batch reinforcement learning (RL) aims at leveraging pre-collected data to find an optimal policy that maximizes the expected total rewards in a dynamic environment. The existing methods require absolutely continuous assumption (e.g., there…

机器学习 · 统计学 2024-06-27 Xiaohong Chen , Zhengling Qi , Runzhe Wan

Estimation of solution norms and stability for time-dependent nonlinear systems is ubiquitous in numerous engineering, natural science and control problems. Yet, practically valuable results are rare in this area. This paper develops a…

动力系统 · 数学 2020-01-22 Mark A. Pinsky , Steve Koblik

Designing model-free algorithms for distributionally robust reinforcement learning (DRRL) poses fundamental challenges. The robust Bellman operator is nonlinear in the transition kernel, which makes one-sample Bellman updates biased, while…

机器学习 · 计算机科学 2026-05-12 Shengbo Wang , Zexi Zhang

Hamilton-Jacobi reachability (HJR) provides a value function that encodes the set of states from which a system with bounded control inputs can reach or avoid a target despite any bounded disturbance, and the corresponding robust, optimal…

系统与控制 · 电气工程与系统科学 2025-06-23 Will Sharpless , Yat Tin Chow , Sylvia Herbert

Representation learning (RL) methods learn objects' latent embeddings where information is preserved by distances. Since distances are invariant to certain linear transformations, one may obtain different embeddings while preserving the…

机器学习 · 计算机科学 2021-01-19 Furkan Gürsoy , Mounir Haddad , Cécile Bothorel

In this work, we propose a methodology for the expression of necessary and sufficient Lyapunov-like conditions for the existence of stabilizing feedback laws. The methodology is an extension of the well-known Control Lyapunov Function (CLF)…

最优化与控制 · 数学 2008-01-31 Iasson Karafyllis , Zhong-Ping Jiang

Deep reinforcement learning has been recognized as a promising tool to address the challenges in real-time control of power systems. However, its deployment in real-world power systems has been hindered by a lack of explicit stability and…

系统与控制 · 电气工程与系统科学 2023-10-04 Jie Feng , Yuanyuan Shi , Guannan Qu , Steven H. Low , Anima Anandkumar , Adam Wierman

This paper develops a model-based reinforcement learning (MBRL) framework for learning online the value function of an infinite-horizon optimal control problem while obeying safety constraints expressed as control barrier functions (CBFs).…

机器学习 · 计算机科学 2022-11-10 Max H. Cohen , Calin Belta

In this paper we apply the formalism of translation invariant (continuous) matrix product states in the thermodynamic limit to $(1+1)$ dimensional critical models. Finite bond dimension bounds the entanglement entropy and introduces an…

量子物理 · 物理学 2015-06-18 Vid Stojevic , Jutho Haegeman , I. P. McCulloch , L. Tagliacozzo , Frank Verstraete

Two main challenges in Reinforcement Learning (RL) are designing appropriate reward functions and ensuring the safety of the learned policy. To address these challenges, we present a theoretical framework for Inverse Reinforcement Learning…

机器学习 · 计算机科学 2023-06-02 Andreas Schlaginhaufen , Maryam Kamgarpour

We study integral-to-integral input-to-state stability for infinite-dimensional linear systems with inputs and trajectories in $L^p$-spaces. We start by developing the corresponding admissibility theory for linear systems with unbounded…

最优化与控制 · 数学 2026-05-26 Sahiba Arora , Andrii Mironchenko

In this article we are interested in the boundary stabilization in finite time of one-dimensional linear hyperbolic balance laws with coefficients depending on time and space. We extend the so called "backstepping method" by introducing…

最优化与控制 · 数学 2020-11-30 Jean-Michel Coron , Long Hu , Guillaume Olive , Peipei Shang

A differential dynamic programming (DDP)-based framework for inverse reinforcement learning (IRL) is introduced to recover the parameters in the cost function, system dynamics, and constraints from demonstrations. Different from existing…

机器人学 · 计算机科学 2024-07-30 Kun Cao , Xinhang Xu , Wanxin Jin , Karl H. Johansson , Lihua Xie

The field of risk-constrained reinforcement learning (RCRL) has been developed to effectively reduce the likelihood of worst-case scenarios by explicitly handling risk-measure-based constraints. However, the nonlinearity of risk measures…

机器学习 · 计算机科学 2024-05-30 Dohyeong Kim , Taehyun Cho , Seungyub Han , Hojun Chung , Kyungjae Lee , Songhwai Oh

Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to settings such as verifiable rewards and multi-objective alignment, RMs are expected to encode…

机器学习 · 计算机科学 2026-05-21 Jiwoo Hong , Shao Tang , Zhipeng Wang

We introduce a novel framework for analyzing reinforcement learning (RL) in continuous state-action spaces, and use it to prove fast rates of convergence in both off-line and on-line settings. Our analysis highlights two key stability…

机器学习 · 计算机科学 2024-01-11 Yaqi Duan , Martin J. Wainwright

We study inventory control policies for pharmaceutical supply chains, addressing challenges such as perishability, yield uncertainty, and non-stationary demand, combined with batching constraints, lead times, and lost sales. Collaborating…

人工智能 · 计算机科学 2025-01-22 Francesco Stranieri , Chaaben Kouki , Willem van Jaarsveld , Fabio Stella