中文
相关论文

相关论文: KL-learning: Online solution of Kullback-Leibler c…

200 篇论文

We study the problem of designing a state feedback linear quadratic Gaussian (LQG) controller for a system in which the system matrices as well as the process noise covariance are unknown. We do a rigorous comparison between two approaches.…

系统与控制 · 电气工程与系统科学 2025-11-13 Mingxiang Liu , Damián Marelli , Minyue Fu , Qianqian Cai

In reinforcement learning (RL), offline learning decoupled learning from data collection and is useful in dealing with exploration-exploitation tradeoff and enables data reuse in many applications. In this work, we study two offline…

机器学习 · 计算机科学 2022-02-08 Jing Dong , Xin T. Tong

This paper is concerned with the linear quadratic optimal control of discrete-time time-varying system with terminal state constraint. The main contribution is to propose a Q-learning algorithm for the optimal controller when the…

最优化与控制 · 数学 2023-07-20 Juanjuan Xu , Jingmei Liu , Zhaorong Zhang , Wei Wang

We quantify the performance of approximations to stochastic filtering by the Kullback-Leibler divergence to the optimal Bayesian filter. Using a two-state Markov process that drives a Brownian measurement process as prototypical test case,…

统计力学 · 物理学 2022-10-26 Rahul O. Ramakrishnan , Andrea Auconi , Benjamin M. Friedrich

The ODE method has been a workhorse for algorithm design and analysis since the introduction of the stochastic approximation. It is now understood that convergence theory amounts to establishing robustness of Euler approximations for ODEs,…

最优化与控制 · 数学 2020-10-02 Shuhang Chen , Adithya Devraj , Andrey Bernstein , Sean Meyn

In this paper we make a survey on the so called randomization method, a recent methodology to study stochastic optimization problems. It allows to represent the value function of an optimal control problem by a suitable backward stochastic…

最优化与控制 · 数学 2025-06-12 Marco Fuhrman

Because reinforcement learning suffers from a lack of scalability, online value (and Q-) function approximation has received increasing interest this last decade. This contribution introduces a novel approximation scheme, namely the Kalman…

机器学习 · 计算机科学 2014-06-13 Matthieu Geist , Olivier Pietquin

This paper presents a distributed learning model predictive control (DLMPC) scheme for distributed linear time invariant systems with coupled dynamics and state constraints. The proposed solution method is based on an online distributed…

系统与控制 · 电气工程与系统科学 2020-06-25 Yvonne R. Stürz , Edward L. Zhu , Ugo Rosolia , Karl H. Johansson , Francesco Borrelli

This paper addresses the problem of learning control policies for mobile robots, modeled as unknown Markov Decision Processes (MDPs), that are tasked with temporal logic missions, such as sequencing, coverage, or surveillance. The MDP…

机器人学 · 计算机科学 2022-07-13 Yiannis Kantaros

This paper presents a novel methodology to tackle feedback optimal control problems in scenarios where the exact state of the controlled process is unknown. It integrates data assimilation techniques and optimal control solvers to manage…

最优化与控制 · 数学 2024-04-10 Siming Liang , Ruoyu Hu , Feng Bao , Richard Archibald , Guannan Zhang

We study the non-contextual multi-armed bandit problem in a transfer learning setting: before any pulls, the learner is given N'_k i.i.d. samples from each source distribution nu'_k, and the true target distributions nu_k lie within a known…

机器学习 · 计算机科学 2025-09-24 Adrien Prevost , Timothee Mathieu , Odalric-Ambrym Maillard

Kernel embeddings of distributions have recently gained significant attention in the machine learning community as a data-driven technique for representing probability distributions. Broadly, these techniques enable efficient computation of…

最优化与控制 · 数学 2021-03-25 Adam J. Thorpe , Meeko M. K. Oishi

We consider the classic stochastic linear quadratic regulator (LQR) problem under an infinite horizon average stage cost. By leveraging recent policy gradient methods from reinforcement learning, we obtain a first-order method that finds a…

最优化与控制 · 数学 2025-02-21 Caleb Ju , Georgios Kotsalis , Guanghui Lan

This article presents a constrained policy optimization approach for the optimal control of systems under nonstationary uncertainties. We introduce an assumption that we call Markov embeddability that allows us to cast the stochastic…

最优化与控制 · 数学 2026-05-11 Sungho Shin , François Pacaud , Emil Contantinescu , Mihai Anitescu

In this study, we introduce numerical methods for discretizing continuous-time linear-quadratic optimal control problems (LQ-OCPs). The discretization of continuous-time LQ-OCPs is formulated into differential equation systems, and we can…

This paper proposes a learning algorithm to find a scheduling policy that achieves an optimal delay-power trade-off in communication systems. Reinforcement learning (RL) is used to minimize the expected latency for a given energy constraint…

系统与控制 · 电气工程与系统科学 2020-06-11 Yu Zhao , Joohyun Lee , Wei Chen

We study online reinforcement learning for finite-horizon deterministic control systems with {\it arbitrary} state and action spaces. Suppose that the transition dynamics and reward function is unknown, but the state and action space is…

机器学习 · 计算机科学 2019-05-07 Lin F. Yang , Chengzhuo Ni , Mengdi Wang

This paper considers the problem of minimizing a convex expectation function with a set of inequality convex expectation constraints. We present a computable stochastic approximation type algorithm, namely the stochastic linearized proximal…

最优化与控制 · 数学 2022-06-16 Liwei Zhang , Yule Zhang , Jia Wu , Xiantao Xiao

This paper illustrates novel methods for nonstationary time series modeling along with their applications to selected problems in neuroscience. These methods are semi-parametric in that inferences are derived by combining sequential…

应用统计 · 统计学 2010-11-03 Fabio Rigat , Jim Q. Smith

We propose an algorithm for approximating the solution of a strongly oscillating SDE, that is, a system in which some ergodic state variables evolve quickly with respect to the other variables. The algorithm profits from homogenization…

概率论 · 数学 2015-03-19 Camilo Andrés García Trillos