中文
相关论文

相关论文: q-exponential family for policy optimization

200 篇论文

Front propagation into unstable states is often determined by the linearization, that is, propagation speeds agree with predictions from the linearized equation at the unstable state. The leading edge behavior is then a Gaussian tail…

偏微分方程分析 · 数学 2025-08-21 Montie Avery , Matt Holzer , Arnd Scheel

We propose the notion of sub-Weibull distributions, which are characterised by tails lighter than (or equally light as) the right tail of a Weibull distribution. This novel class generalises the sub-Gaussian and sub-Exponential families to…

统计理论 · 数学 2020-12-04 Mariia Vladimirova , Stephane Girard , Hien Nguyen , Julyan Arbel

For purposes of Value-at-Risk estimation, we consider several multivariate families of heavy-tailed distributions, which can be seen as multidimensional versions of Paretian stable and Student's t distributions allowing different marginals…

风险管理 · 定量金融 2011-12-20 Carlo Marinelli , Stefano d'Addona , Svetlozar T. Rachev

In this work, we introduce a Z-control strategy for multi-agent systems of arbitrary order, aimed at driving the agents toward consensus in the highest-order observable state. The proposed framework supports both direct and indirect control…

最优化与控制 · 数学 2025-11-26 Angela Monti , Fasma Diele

In this technical note, we establish an upper-bound on the threshold on the discount factor starting from which all discounted-optimal deterministic policies are gain-optimal, that we prove to be tight on an example. To address…

系统与控制 · 电气工程与系统科学 2023-04-18 Victor Boone

The widespread diffusion of mobile phones is triggering an exponential growth of mobile data traffic that is likely to cause, in the near future, considerable traffic overload issues even in last-generation cellular networks. Offloading…

网络与互联网体系结构 · 计算机科学 2021-10-04 Lorenzo Valerio , Raffaele Bruno , Andrea Passarella

This work deals with the time evolution of the Hamming distance density for the public goods game. We consider distinct possibilities for this game, which are exactly described by a function called $q$-exponential, that represents a…

物理与社会 · 物理学 2024-03-29 D. Bazeia , M. J. B. Ferreira , B. F. de Oliveira

On-policy reinforcement learning (RL) algorithms are widely used for their strong asymptotic performance and training stability, but they struggle to scale with larger batch sizes, as additional parallel environments yield redundant data…

机器学习 · 计算机科学 2025-11-13 Jianren Wang , Yifan Su , Abhinav Gupta , Deepak Pathak

Policy gradient methods are reinforcement learning algorithms that adapt a parameterized policy by following a performance gradient estimate. Conventional policy gradient methods use Monte-Carlo techniques to estimate the gradient, which…

机器学习 · 计算机科学 2026-05-01 Mohammad Ghavamzadeh , Yaakov Engel , Michal Valko

We consider reinforcement learning (RL) methods in offline domains without additional online data collection, such as mobile health applications. Most of existing policy optimization algorithms in the computer science literature are…

机器学习 · 统计学 2022-07-28 Chengchun Shi , Shikai Luo , Yuan Le , Hongtu Zhu , Rui Song

One of the key performance measures in queueing systems is the exponential decay rate of the steady-state tail probabilities of the queue lengths. It is known that if a corresponding fluid model is stable and the stochastic primitives have…

概率论 · 数学 2007-05-23 David Gamarnik , Sean Meyn

We prove that output-feedback linear policies remain optimal for solving the Linear Quadratic Gaussian regulation problem in the face of worst-case process and measurement noise distributions when these are independent, stationary, and…

最优化与控制 · 数学 2025-04-23 Nicolas Lanzetti , Antonio Terpin , Florian Dörfler

We describe a novel approach to explainable prediction of a continuous variable based on learning fuzzy weighted rules. Our model trains a set of weighted rules to maximise prediction accuracy and minimise an ontology-based 'semantic loss'…

人工智能 · 计算机科学 2022-08-29 Martin Glauer , Robert West , Susan Michie , Janna Hastings

This paper considers linear quadratic team decision problems where the players in the team affect each other's information structure through their decisions. Whereas the stochastic version of the problem is well known to be complex with…

最优化与控制 · 数学 2013-02-05 Ather Gattami

We propose Q-learning with Adjoint Matching (QAM), a novel TD-based reinforcement learning (RL) algorithm that tackles a long-standing challenge in continuous-action RL: efficient optimization of an expressive diffusion or flow-matching…

机器学习 · 计算机科学 2026-05-20 Qiyang Li , Sergey Levine

We propose Q-Policy, a hybrid quantum-classical reinforcement learning (RL) framework that mathematically accelerates policy evaluation and optimization by exploiting quantum computing primitives. Q-Policy encodes value functions in quantum…

机器学习 · 计算机科学 2025-06-10 Kalyan Cherukuri , Aarav Lala , Yash Yardi

In an online decision problem, one makes decisions often with a pool of decision sequence called experts but without knowledge of the future. After each step, one pays a cost based on the decision and observed rate. One reasonal goal would…

机器学习 · 计算机科学 2015-12-23 Chunyang Xiao

Learning methods are increasingly used to synthesize controllers from data, yet existing sample-complexity characterizations for continuous control are sharp only in the fully observed setting. This paper studies the partially observed case…

系统与控制 · 电气工程与系统科学 2026-05-19 Bruce D. Lee , Anastasios Tsiamis , Nikolai Matni , Manfred Morari , John Lygeros

We consider a learning system based on the conventional multiplicative weight (MW) rule that combines experts' advice to predict a sequence of true outcomes. It is assumed that one of the experts is malicious and aims to impose the maximum…

机器学习 · 计算机科学 2020-09-21 S. Rasoul Etesami , Negar Kiyavash , Vincent Leon , H. Vincent Poor

Guided exploration with expert demonstrations improves data efficiency for reinforcement learning, but current algorithms often overuse expert information. We propose a novel algorithm to speed up Q-learning with the help of a limited…

机器学习 · 计算机科学 2022-10-06 Fengdi Che , Xiru Zhu , Doina Precup , David Meger , Gregory Dudek