中文
相关论文

相关论文: Single Time-scale Actor-critic Method to Solve the…

200 篇论文

This paper addresses a risk-constrained decentralized stochastic linear-quadratic optimal control problem with one remote controller and one local controller, where the risk constraint is posed on the cumulative state weighted variance in…

最优化与控制 · 数学 2023-07-19 Jia Hui , Yuan-Hua Ni

This paper studies a class of continuous-time scalar-state stochastic Linear-Quadratic (LQ) optimal control problem with the linear control constraints. Applying the state separation theorem induced from its special structure, we develop…

投资组合管理 · 定量金融 2018-06-12 Weiping Wu , Jianjun Gao , Junguo Lu , Xun Li

We propose a new risk-constrained reformulation of the standard Linear Quadratic Regulator (LQR) problem. Our framework is motivated by the fact that the classical (risk-neutral) LQR controller, although optimal in expectation, might be…

系统与控制 · 电气工程与系统科学 2020-10-30 Anastasios Tsiamis , Dionysios S. Kalogerias , Luiz F. O. Chamon , Alejandro Ribeiro , George J. Pappas

A fundamental theory of deterministic linear-quadratic (LQ) control is the equivalent relationship between control problems, two-point boundary value problems and Riccati equations. In this paper, we extend the equivalence to a general…

数理金融 · 定量金融 2021-10-13 Hongyan Cai , Danhong Chen , Yunfei Peng , Wei Wei

Explicit solutions to optimal control problems are rarely obtainable. Of particular interest are the explicit solutions derived for minimax problems, providing a framework to address adversarial conditions and uncertainty. This work…

最优化与控制 · 数学 2026-03-10 Alba Gurpegui , Mark Jeeninga , Emma Tegling , Anders Rantzer

A time-inconsistent optimal control problem is formulated and studied for a controlled linear ordinary differential equation with quadratic cost functional. A notion of equilibrium control is introduced, which can be regarded as a…

最优化与控制 · 数学 2012-04-10 Jiongmin Yong

Classical linear quadratic (LQ) control centers around linear time-invariant (LTI) systems, where the control-state pairs introduce a quadratic cost with time-invariant parameters. Recent advancement in online optimization and control has…

最优化与控制 · 数学 2020-09-30 Ting-Jui Chang , Shahin Shahrampour

We study the sample complexity of approximate policy iteration (PI) for the Linear Quadratic Regulator (LQR), building on a recent line of work using LQR as a testbed to understand the limits of reinforcement learning (RL) algorithms on…

机器学习 · 计算机科学 2019-05-31 Karl Krauth , Stephen Tu , Benjamin Recht

Despite the popularity of the actor-critic method and the practical needs of collaborative policy training, existing works typically either overlook environmental heterogeneity or give up personalization altogether by training a single…

机器学习 · 计算机科学 2026-05-15 Leo Muxing Wang , Pengkun Yang , Lili Su

This paper studies data-driven approaches to the continuous-time linear quadratic regulator (LQR) problem based on two existing parameterizations, namely a closed-loop (CL) parameterization from behavioral system theory and an integral…

最优化与控制 · 数学 2026-05-01 Armin Gießler , Felix Thömmes , Sören Hohmann

This paper presents a sample-efficient, data-driven control framework for finite-horizon linear quadratic (LQ) control of linear time-varying (LTV) systems. In contrast to the time-invariant case, the time-varying LQ problem involves a…

系统与控制 · 电气工程与系统科学 2025-09-30 Sahel Vahedi Noori , Maryam Babazadeh

Linear-Quadratic (LQ) problems that arise in systems and controls include the classical optimal control problems of the Linear Quadratic Regulator (LQR) in both its deterministic and stochastic forms, as well as $H^\infty$-analysis (the…

系统与控制 · 电气工程与系统科学 2024-01-04 Bassam Bamieh

Actor-critic (AC) methods are widely used in reinforcement learning (RL) and benefit from the flexibility of using any policy gradient method as the actor and value-based method as the critic. The critic is usually trained by minimizing the…

机器学习 · 计算机科学 2023-11-01 Sharan Vaswani , Amirreza Kazemi , Reza Babanezhad , Nicolas Le Roux

In this paper, a cooperative Linear Quadratic Regulator (LQR) problem is investigated for multi-input systems, where each input is generated by an agent in a network. The input matrices are different and locally possessed by the…

多智能体系统 · 计算机科学 2021-11-10 Peihu Duan , Lidong He , Zhisheng Duan , Ling Shi

Linear-quadratic optimal control problems are considered for mean-field stochastic differential equations with deterministic coefficients. Time-inconsistency feature of the problems is carefully investigated. Both open-loop and closed-loop…

最优化与控制 · 数学 2013-05-07 Jiongmin Yong

We propose a stochastic approximation (SA) based method with randomization of samples for policy evaluation using the least squares temporal difference (LSTD) algorithm. Our proposed scheme is equivalent to running regular temporal…

机器学习 · 计算机科学 2020-01-27 L. A. Prashanth , Nathaniel Korda , Rémi Munos

This work contributes to the field of optimal control of bilinear systems. It concerns a continuous time, finite dimensional, bilinear state equation with a quadratic performance index to be minimized. The state equation is non-autonomous…

最优化与控制 · 数学 2022-05-02 Ido Halperin

This paper addresses the optimal control problem known as the Linear Quadratic Regulator in the case when the dynamics are unknown. We propose a multi-stage procedure, called Coarse-ID control, that estimates a model from a few experimental…

最优化与控制 · 数学 2018-12-17 Sarah Dean , Horia Mania , Nikolai Matni , Benjamin Recht , Stephen Tu

We propose and analyse a new methodology based on linear-quadratic regulation (LQR) for stabilising falling liquid films via blowing and suction at the base. LQR methods enable rapidly responding feedback control by precomputing a gain…

最优化与控制 · 数学 2023-07-11 Oscar A. Holroyd , Radu Cimpeanu , Susana N. Gomes

We revisit the standard formulation of tabular actor-critic algorithm as a two time-scale stochastic approximation with value function computed on a faster time-scale and policy computed on a slower time-scale. This emulates policy…

机器学习 · 计算机科学 2024-06-21 Shalabh Bhatnagar , Vivek S. Borkar , Soumyajit Guin