English
Related papers

Related papers: KL-learning: Online solution of Kullback-Leibler c…

200 papers

We consider an offline reinforcement learning (RL) setting where the agent need to learn from a dataset collected by rolling out multiple behavior policies. There are two challenges for this setting: 1) The optimal trade-off between…

Machine Learning · Statistics 2022-12-06 Yuanying Cai , Chuheng Zhang , Li Zhao , Wei Shen , Xuyun Zhang , Lei Song , Jiang Bian , Tao Qin , Tieyan Liu

We study a variant of online optimization in which the learner receives $k$-round $\textit{delayed feedback}$ about hitting cost and there is a multi-step nonlinear switching cost, i.e., costs depend on multiple previous actions in a…

Machine Learning · Computer Science 2021-11-02 Weici Pan , Guanya Shi , Yiheng Lin , Adam Wierman

In the regret-based formulation of Multi-armed Bandit (MAB) problems, except in rare instances, much of the literature focuses on arms with i.i.d. rewards. In this paper, we consider the problem of obtaining regret guarantees for MAB…

Machine Learning · Computer Science 2022-10-11 Arghyadip Roy , Sanjay Shakkottai , R. Srikant

Deep neural networks are increasingly used as an effective parameterization of control policies in various learning-based control paradigms. For continuous-time optimal control problems (OCPs), which are central to many decision-making…

Machine Learning · Computer Science 2025-11-04 Joshua Hang Sai Ip , Georgios Makrygiorgos , Ali Mesbah

This paper addresses the problem of approximating an unknown probability distribution with density $f$ -- which can only be evaluated up to an unknown scaling factor -- with the help of a sequential algorithm that produces at each iteration…

Statistics Theory · Mathematics 2024-09-23 Pascal Bianchi , Bernard Delyon , Victor Priser , François Portier

The paper investigates stochastic resource allocation problems with scarce, reusable resources and non-preemtive, time-dependent, interconnected tasks. This approach is a natural generalization of several standard resource management…

Machine Learning · Computer Science 2014-01-16 Balázs Csanád Csáji , László Monostori

We study an extension of the classic stochastic multi-armed bandit problem which involves multiple plays and Markovian rewards in the rested bandits setting. In order to tackle this problem we consider an adaptive allocation rule which at…

Statistics Theory · Mathematics 2020-07-15 Vrettos Moulos

Distributionally Robust Optimization (DRO), as a popular method to train robust models against distribution shift between training and test sets, has received tremendous attention in recent years. In this paper, we propose and analyze…

Machine Learning · Computer Science 2023-08-17 Qi Qi , Jiameng Lyu , Kung sik Chan , Er Wei Bai , Tianbao Yang

The Kullback-Leibler (KL) divergence is a fundamental equation of information theory that quantifies the proximity of two probability distributions. Although difficult to understand by examining the equation, an intuition and understanding…

Information Theory · Computer Science 2014-04-09 Jonathon Shlens

Recent years have seen a growing interest in understanding acceleration methods through the lens of ordinary differential equations (ODEs). Despite the theoretical advancements, translating the rapid convergence observed in continuous-time…

Optimization and Control · Mathematics 2024-06-05 Zhonglin Xie , Wotao Yin , Zaiwen Wen

In this paper the connection between stochastic optimal control and reinforcement learning is investigated. Our main motivation is to apply importance sampling to sampling rare events which can be reformulated as an optimal control problem.…

Optimization and Control · Mathematics 2024-02-16 Jannes Quer , Enric Ribera Borrell

The LogSumExp function, dual to the Kullback-Leibler (KL) divergence, plays a central role in many important optimization problems, including entropy-regularized optimal transport (OT) and distributionally robust optimization (DRO). In…

Optimization and Control · Mathematics 2026-02-04 Egor Gladin , Alexey Kroshnin , Jia-Jie Zhu , Pavel Dvurechensky

The authors consider stochastic aspects of the stabilization problem for two and three-dimensional Oseen equations with help of feedback control defined on a part of the fluid boundary. Stochastic issues arise when inevitable unpredictable…

Analysis of PDEs · Mathematics 2007-05-23 Jinqiao Duan , Andrei V. Fursikov

To overcome the curses of dimensionality and modeling of Dynamic Programming (DP) methods to solve Markov Decision Process (MDP) problems, Reinforcement Learning (RL) methods are adopted in practice. Contrary to traditional RL algorithms…

Machine Learning · Computer Science 2021-08-24 Arghyadip Roy , Vivek Borkar , Abhay Karandikar , Prasanna Chaporkar

This paper studies the linear quadratic regulation (LQR) problem of unknown discrete-time systems via dynamic output feedback learning control. In contrast to the state feedback, the optimality of the dynamic output feedback control for…

Systems and Control · Electrical Eng. & Systems 2025-05-29 Kedi Xie , Martin Guay , Shimin Wang , Fang Deng , Maobin Lu

In this paper, we consider a distributionally robust optimization (DRO) model in which the ambiguity set is defined as the set of distributions whose Kullback-Leibler (KL) divergence to an empirical distribution is bounded. Utilizing the…

Optimization and Control · Mathematics 2024-11-12 Burak Kocuk

We study the problem of online learning in predictive control of an unknown linear dynamical system with time varying cost functions which are unknown apriori. Specifically, we study the online learning problem where the control algorithm…

Machine Learning · Computer Science 2022-11-01 Deepan Muthirayan , Jianjun Yuan , Dileep Kalathil , Pramod P. Khargonekar

We recently proposed a general algorithm for approximating nonstandard Bayesian posterior distributions by minimization of their Kullback-Leibler divergence with respect to a more convenient approximating distribution. In this note we offer…

Computation · Statistics 2014-01-10 Tim Salimans

We present a model-free reinforcement learning algorithm to find an optimal policy for a finite-horizon Markov decision process while guaranteeing a desired lower bound on the probability of satisfying a signal temporal logic (STL)…

Systems and Control · Electrical Eng. & Systems 2021-09-29 Krishna C. Kalagarla , Rahul Jain , Pierluigi Nuzzo

We study a type of Online Linear Programming (OLP) problem that maximizes the objective function with stochastic inputs. The performance of various algorithms that analyze this type of OLP is well studied when the stochastic inputs follow…

Optimization and Control · Mathematics 2022-10-04 Owen Shen