中文
相关论文

相关论文: KL-learning: Online solution of Kullback-Leibler c…

200 篇论文

A stochastic combinatorial semi-bandit is an online learning problem where at each step a learning agent chooses a subset of ground items subject to constraints, and then observes stochastic weights of these items and receives their sum as…

机器学习 · 计算机科学 2017-06-08 Branislav Kveton , Zheng Wen , Azin Ashkan , Csaba Szepesvari

The solution to a stochastic optimal control problem can be determined by computing the value function from a discretization of the associated Hamilton-Jacobi-Bellman equation. Alternatively, the problem can be reformulated in terms of a…

最优化与控制 · 数学 2024-02-29 Sebastian Reich

In this paper, we propose a new policy iteration algorithm to compute the value function and the optimal controls of continuous time stochastic control problems. The algorithm relies on successive approximations using linear-quadratic…

最优化与控制 · 数学 2024-09-09 Dylan Possamaï , Ludovic Tangpi

PDE-constrained optimization problems arise in a broad number of applications such as hyperthermia cancer treatment or blood flow simulation. Discretization of the optimization problem and using a Lagrangian approach result in a large-scale…

数值分析 · 数学 2020-06-01 Alexandra Bünger , Valeria Simoncini , Martin Stoll

In this paper long-run risk sensitive optimisation problem is studied with dyadic impulse control applied to continuous-time Feller-Markov process. In contrast to the existing literature, focus is put on unbounded and non-uniformly ergodic…

最优化与控制 · 数学 2019-06-18 Marcin Pitera , Łukasz Stettner

This paper focuses on adaptive control of the discrete-time linear quadratic regulator (adaptive LQR). Recent literature has made significant contributions in proving non-asymptotic convergence rates, but existing approaches have a few…

系统与控制 · 电气工程与系统科学 2026-04-27 Peter A. Fisher , Anuradha M. Annaswamy

We propose a novel approach to solve K-adaptability problems with convex objective and constraints and integer first-stage decisions. A logic-based Benders decomposition is applied to handle the first-stage decisions in a master problem,…

最优化与控制 · 数学 2022-09-08 Alireza Ghahtarani , Ahmed Saif , Alireza Ghasemi , Erick Delage

In this paper, we consider the problem of real-time transmission scheduling over time-varying channels. We first formulate the transmission scheduling problem as a Markov decision process (MDP) and systematically unravel the structural…

机器学习 · 计算机科学 2010-03-15 Fangwen Fu , Mihaela van der Schaar

We study a general class of PageRank optimization problems which consist in finding an optimal outlink strategy for a web site subject to design constraints. We consider both a continuous problem, in which one can choose the intensity of a…

最优化与控制 · 数学 2016-01-08 Olivier Fercoq , Marianne Akian , Mustapha Bouhtou , Stéphane Gaubert

We study regret minimization in online episodic linear Markov Decision Processes, and obtain rate-optimal $\widetilde O (\sqrt K)$ regret where $K$ denotes the number of episodes. Our work is the first to establish the optimal (w.r.t.~$K$)…

机器学习 · 计算机科学 2024-05-17 Uri Sherman , Alon Cohen , Tomer Koren , Yishay Mansour

In this contribution, we derive ILEG, an iterative algorithm to find risk sensitive solutions to nonlinear, stochastic optimal control problems. The algorithm is based on a linear quadratic approximation of an exponential risk sensitive…

系统与控制 · 计算机科学 2015-12-23 Farbod Farshidian , Jonas Buchli

Real-world autonomous systems operate under uncertainty about both their pose and dynamics. Autonomous control systems must simultaneously perform estimation and control tasks to maintain robustness to changing dynamics or modeling errors.…

系统与控制 · 计算机科学 2018-08-03 Patrick Slade , Zachary N. Sunberg , Mykel J. Kochenderfer

Data collection is a critical step in statistical inference and data science, and the goal of statistical experimental design (ED) is to find the data collection setup that can provide most information for the inference. In this work we…

统计计算 · 统计学 2020-07-01 Ziqiao Ao , Jinglai Li

This paper is concerned with a linear-quadratic (LQ, for short) optimal control problem for backward stochastic differential equations (BSDEs, for short), where the coefficients of the backward control system and the weighting matrices in…

最优化与控制 · 数学 2021-05-14 Jingrui Sun , Hanxiao Wang

This paper develops KL-Ergodic Exploration from Equilibrium ($\text{KL-E}^3$), a method for robotic systems to integrate stability into actively generating informative measurements through ergodic exploration. Ergodic exploration enables…

机器人学 · 计算机科学 2020-12-08 Ian Abraham , Ahalya Prabhakar , Todd D. Murphey

This paper is concerned with a stochastic linear quadratic (LQ, for short) control problem with a recursive cost functional in an infinite horizon. A main difficult is well-posedness of the BSDE in $L^1$ and in infinite horizon. A notion of…

最优化与控制 · 数学 2026-05-07 Lin Li , Jiongmin Yong

We study the policy gradient method (PGM) for the linear quadratic Gaussian (LQG) dynamic output-feedback control problem using an input-output-history (IOH) representation of the closed-loop system. First, we show that any dynamic…

最优化与控制 · 数学 2025-10-23 Tomonori Sadamoto , Takashi Tanaka

Consider the problem of approximating the optimal policy of a Markov decision process (MDP) by sampling state transitions. In contrast to existing reinforcement learning methods that are based on successive approximations to the nonlinear…

机器学习 · 计算机科学 2017-10-18 Mengdi Wang

We study the problem of zero-delay coding for the transmission of a Markov source over a noisy channel with feedback and present a reinforcement learning solution which is guaranteed to achieve near-optimality. To this end, we formulate the…

最优化与控制 · 数学 2025-10-07 Liam Cregg , Fady Alajaji , Serdar Yuksel

We study the inverse optimal control problem in social sciences: we aim at learning a user's true cost function from the observed temporal behavior. In contrast to traditional phenomenological works that aim to learn a generative model to…

机器学习 · 计算机科学 2018-05-23 Yichen Wang , Le Song , Hongyuan Zha