中文
相关论文

相关论文: On the relation between dynamic regret and closed-…

200 篇论文

Predictive safety filters provide a way of projecting potentially unsafe inputs, proposed, e.g. by a human or learning-based controller, onto the set of inputs that guarantee recursive state and input constraint satisfaction by leveraging…

系统与控制 · 电气工程与系统科学 2024-04-30 Alexandre Didier , Andrea Zanelli , Kim P. Wabersich , Melanie N. Zeilinger

We study the non-stationary stochastic multi-armed bandit problem, where the reward statistics of each arm may change several times during the course of learning. The performance of a learning algorithm is evaluated in terms of their…

机器学习 · 计算机科学 2022-03-09 Yasin Abbasi-Yadkori , Andras Gyorgy , Nevena Lazic

We introduce algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary linear stochastic bandit setting. It captures natural applications such as dynamic pricing and ads allocation in a changing environment.…

机器学习 · 计算机科学 2021-07-20 Wang Chi Cheung , David Simchi-Levi , Ruihao Zhu

This paper presents a class of Dynamic Multi-Armed Bandit problems where the reward can be modeled as the noisy output of a time varying linear stochastic dynamic system that satisfies some boundedness constraints. The class allows many…

机器学习 · 计算机科学 2017-10-10 T. W. U. Madhushani , D. H. S. Maithripala , N. E. Leonard

Model Predictive Control (MPC) is well understood in the deterministic setting, yet rigorous stability and performance guarantees for stochastic MPC remain limited to the consideration of terminal constraints and penalties. In contrast,…

This work focuses on the setting of dynamic regret in the context of online learning with full information. In particular, we analyze regret bounds with respect to the temporal variability of the loss functions. By assuming that the…

机器学习 · 计算机科学 2021-02-16 Nicolò Campolongo , Francesco Orabona

We study the existence of asymptotically stable periodic trajectories induced by reset feedback. The analysis is developed for a planar system. Casting the problem into the hybrid setting, we show that a periodic orbit arises from the…

系统与控制 · 计算机科学 2021-02-26 Andrea Bisoffi , Fulvio Forni , Mauro Da Lio , Luca Zaccarian

We study stochastic structured bandits for minimizing regret. The fact that the popular optimistic algorithms do not achieve the asymptotic instance-dependent regret optimality (asymptotic optimality for short) has recently alluded…

机器学习 · 计算机科学 2020-10-26 Kwang-Sung Jun , Chicheng Zhang

We investigate online Markov Decision Processes (MDPs) with adversarially changing loss functions and known transitions. We choose dynamic regret as the performance measure, defined as the performance difference between the learner and any…

机器学习 · 计算机科学 2022-08-29 Peng Zhao , Long-Fei Li , Zhi-Hua Zhou

We study the problem of \emph{dynamic regret minimization} in $K$-armed Dueling Bandits under non-stationary or time varying preferences. This is an online learning setup where the agent chooses a pair of items at each round and observes…

机器学习 · 计算机科学 2022-06-14 Aadirupa Saha , Shubham Gupta

Our aim in this paper is to investigate the asymptotic behavior of solutions of the perturbed linear fractional differential system. We show that if the original linear autonomous system is asymptotically stable then under the action of…

动力系统 · 数学 2018-08-24 N. D. Cong , T. S. Doan , H. T. Tuan

We derive a novel asymptotic problem-dependent lower-bound for regret minimization in finite-horizon tabular Markov Decision Processes (MDPs). While, similar to prior work (e.g., for ergodic MDPs), the lower-bound is the solution to an…

机器学习 · 计算机科学 2021-06-25 Andrea Tirinzoni , Matteo Pirotta , Alessandro Lazaric

We consider the problem of online control of systems with time-varying linear dynamics. This is a general formulation that is motivated by the use of local linearization in control of nonlinear dynamical systems. To state meaningful…

机器学习 · 计算机科学 2022-02-15 Paula Gradu , Elad Hazan , Edgar Minasyan

Adaptively controlling and minimizing regret in unknown dynamical systems while controlling the growth of the system state is crucial in real-world applications. In this work, we study the problem of stabilization and regret minimization of…

系统与控制 · 电气工程与系统科学 2022-02-10 Jafar Abbaszadeh Chekan , Kamyar Azizzadenesheli , Cedric Langbort

This work deals with the stability analysis of nonlinear sampled-data systems under nonuniform sampling. It establishes novel relationships between the stability property of the exact discrete-time model for a given sequence of (aperiodic)…

系统与控制 · 电气工程与系统科学 2022-09-28 Alexis J. Vallarella , Hernan Haimovich

It is well-known that the fundamental diagram in a realistic traffic system is featured by capacity drop. From a mesoscopic approach, we demonstrate that such a phenomenon is linked to the unique properties of stochastic noise, which, when…

应用统计 · 统计学 2025-03-21 Mariana Pereira de Melo , Leon Alexander Valencia , Wei-Liang Qian

"Dynamic compensation" is a robustness property where a perturbed biological circuit maintains a suitable output [Karin O., Swisa A., Glaser B., Dor Y., Alon U. (2016). Mol. Syst. Biol., 12: 886]. In spite of several attempts, no fully…

系统与控制 · 计算机科学 2018-01-17 Michel Fliess , Cédric Join

In constrained Markov decision processes (CMDPs) with adversarial rewards and constraints, a well-known impossibility result prevents any algorithm from attaining both sublinear regret and sublinear constraint violation, when competing…

机器学习 · 计算机科学 2024-09-27 Francesco Emanuele Stradi , Anna Lunghi , Matteo Castiglioni , Alberto Marchesi , Nicola Gatti

We introduce data-driven decision-making algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary bandit settings. These settings capture applications such as advertisement allocation, dynamic pricing, and…

机器学习 · 计算机科学 2021-03-19 Wang Chi Cheung , David Simchi-Levi , Ruihao Zhu

The stability of solutions to evolution equations with respect to small stochastic perturbations is considered. The stability of a stochastic dynamical system is characterized by the local stability index. The limit of this index with…

凝聚态物理 · 物理学 2009-11-07 V. I. Yukalov