中文
相关论文

相关论文: Learning Expected Reward for Switched Linear Contr…

200 篇论文

Our goal is to present the basic results on one-dimensional Gibbs and equilibrium states viewed as special invariant measures on symbolic dynamical systems, and then to describe without technicalities a sample of results they allowed to…

动力系统 · 数学 2020-07-16 J. -R. Chazottes , G. Keller

We consider reinforcement learning in changing Markov Decision Processes where both the state-transition probabilities and the reward functions may vary over time. For this problem setting, we propose an algorithm using a sliding window…

机器学习 · 计算机科学 2018-05-28 Pratik Gajane , Ronald Ortner , Peter Auer

We propose a compositional approach to synthesize policies for networks of continuous-space stochastic control systems with unknown dynamics using model-free reinforcement learning (RL). The approach is based on implicitly abstracting each…

系统与控制 · 电气工程与系统科学 2022-08-09 Abolfazl Lavaei , Mateo Perez , Milad Kazemi , Fabio Somenzi , Sadegh Soudjani , Ashutosh Trivedi , Majid Zamani

We present some new results on sample path optimality for the ergodic control problem of a class of non-degenerate diffusions controlled through the drift. The hypothesis most often used in the literature to ensure the existence of an a.s.…

最优化与控制 · 数学 2019-03-20 Ari Arapostathis

This paper investigates gradient-based adaptive prediction and control for nonlinear stochastic dynamical systems under a weak convexity condition on the prediction-based loss. This condition accommodates a broad range of nonlinear models…

系统与控制 · 电气工程与系统科学 2026-02-13 Yujing Liu , Xin Zheng , Zhixin Liu , Lei Guo

Both fixed-gain control and adaptive learning architectures aim to mitigate the effects of uncertainties. In particular, fixed-gain control offers more predictable closed-loop system behavior but requires the knowledge of uncertainty…

系统与控制 · 电气工程与系统科学 2024-03-29 Tansel Yucelen , Selahattin Burak Sarsilmaz , Emre Yildirim

The main challenge for adaptive regulation of linear-quadratic systems is the trade-off between identification and control. An adaptive policy needs to address both the estimation of unknown dynamics parameters (exploration), as well as the…

系统与控制 · 计算机科学 2019-04-01 Mohamad Kazem Shirani Faradonbeh , Ambuj Tewari , George Michailidis

We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under temporal drifts, ie, both the reward and state transition distributions are allowed to evolve over time, as long as their respective total…

机器学习 · 计算机科学 2020-05-19 Wang Chi Cheung , David Simchi-Levi , Ruihao Zhu

Risk-sensitive control balances performance with resilience to unlikely events in uncertain systems. This paper introduces ergodic-risk criteria, which capture long-term cumulative risks through probabilistic limit theorems. By ensuring the…

最优化与控制 · 数学 2025-03-11 Shahriar Talebi , Na Li

Safety in stochastic control systems, which are subject to random noise with a known probability distribution, aims to compute policies that satisfy predefined operational constraints with high confidence throughout the uncertain evolution…

系统与控制 · 电气工程与系统科学 2025-11-12 Saber Omidi , Marek Petrik , Se Young Yoon , Momotaz Begum

The stochastic multi-armed bandit has provided a framework for studying decision-making in unknown environments. We propose a variant of the stochastic multi-armed bandit where the rewards are sampled from a stochastic linear dynamical…

机器学习 · 计算机科学 2022-04-13 Jonathan Gornet , Mehdi Hosseinzadeh , Bruno Sinopoli

Symbiotic control synergistically integrates fixed-gain control and adaptive learning architectures to mitigate system uncertainties more predictably than adaptive learning alone and without requiring prior knowledge of uncertainty bounds…

系统与控制 · 电气工程与系统科学 2024-11-18 Emre Yildirim , Tansel Yucelen , John T. Hrynuk

We present a numerical method for learning unknown nonautonomous stochastic dynamical system, i.e., stochastic system subject to time dependent excitation or control signals. Our basic assumption is that the governing equations for the…

机器学习 · 计算机科学 2025-03-04 Yuan Chen , Dongbin Xiu

We study a stability property of probability laws with respect to small violations of algorithmic randomness. A sufficient condition of stability is presented in terms of Schnorr tests of algorithmic randomness. Most probability laws, like…

计算复杂性 · 计算机科学 2014-09-16 Vladimir V. V'yugin

We study data-driven learning of robust stochastic control for infinite-horizon systems with potentially continuous state and action spaces. In many managerial settings--supply chains, finance, manufacturing, services, and dynamic…

机器学习 · 统计学 2025-11-18 Shengbo Wang , Jason Meng , Nian Si , Jose Blanchet , Zhengyuan Zhou

This paper considers an ergodic version of the bounded velocity follower problem, assuming that the decision maker lacks knowledge of the underlying system parameters and must learn them while simultaneously controlling. We propose…

机器学习 · 统计学 2024-10-07 Stefan Ankirchner , Sören Christensen , Jan Kallsen , Philip Le Borne , Stefan Perko

Linear response theory, the backbone of non-equilibrium statistical physics, has recently been extended to explain how and why non-ergodic renewal processes are insensitive to simple perturbations, such as in habituation. It was established…

适应与自组织系统 · 物理学 2016-06-08 Nicola Piccinini , David Lambert , Bruce West , Mauro Bologna , Paolo Grigolini

We consider the extreme value statistics of centrally-biased random walks with asymptotically-zero drift in the ergodic regime. We fully characterize the asymptotic distribution of the maximum for this class of Markov chains lacking…

统计力学 · 物理学 2022-11-28 Roberto Artuso , Manuele Onofri , Gaia Pozzoli , Mattia Radice

This paper presents a model-free reinforcement learning (RL) algorithm to synthesize a control policy that maximizes the satisfaction probability of linear temporal logic (LTL) specifications. Due to the consideration of environment and…

形式语言与自动机理论 · 计算机科学 2022-01-04 Mingyu Cai , Shaoping Xiao , Baoluo Li , Zhiliang Li , Zhen Kan

We study the estimation of the value function for continuous-time Markov diffusion processes using a single, discretely observed ergodic trajectory. Our work provides non-asymptotic statistical guarantees for the least-squares…

机器学习 · 计算机科学 2025-02-07 Wenlong Mou