中文
相关论文

相关论文: Generator Parameter Estimation by Q-Learning Based…

200 篇论文

This paper presents a framework of learning parameter space for event-triggered control. In particular, our goal is to find a set of parameters for the event-triggered condition, such that certain specifications on safety and convergence…

最优化与控制 · 数学 2020-10-20 Kazumune Hashimoto

In Reinforcement Learning (abbreviated as RL), an agent interacts with the environment via a set of possible actions, and a reward is generated from some unknown distribution. The task here is to find an optimal set of actions such that the…

机器学习 · 计算机科学 2025-07-21 Aditi Anand , Suman Banerjee , Dildar Ali

Adaptive variational algorithms suffer from prohibitively high measurement costs during the generator selection step, since energy gradients must be estimated for a large operator pool. This scaling bottleneck limits their applicability to…

量子物理 · 物理学 2025-09-19 Rick Huang , Artur F. Izmaylov

In this paper, we compare four measures of the empirical observability gramian, including the determinant, the trace, the minimum eigenvalue, and the condition number, which can be used to quantify the observability of system states and to…

最优化与控制 · 数学 2016-06-09 Junjian Qi , Kai Sun , Wei Kang

Histogram-based template fits are the main technique used for estimating parameters of high energy physics Monte Carlo generators. Parametrized neural network reweighting can be used to extend this fitting procedure to many dimensions and…

高能物理 - 唯象学 · 物理学 2021-04-08 Anders Andreassen , Shih-Chieh Hsu , Benjamin Nachman , Natchanon Suaysom , Adi Suresh

The use of machine learning algorithms is an attractive way to produce very fast detector simulations for scattering reactions that can otherwise be computationally expensive. Here we develop a factorised approach where we deal with each…

数据分析、统计与概率 · 物理学 2022-07-26 D. Darulis , R. Tyson , D. G. Ireland , D. I. Glazier , B. McKinnon , P. Pauli

We propose a new simple and natural algorithm for learning the optimal Q-value function of a discounted-cost Markov Decision Process (MDP) when the transition kernels are unknown. Unlike the classical learning algorithms for MDPs, such as…

最优化与控制 · 数学 2019-01-31 Dileep Kalathil , Vivek S. Borkar , Rahul Jain

Approximation errors must be taken into account when compiling quantum programs into a low-level gate set. We present a methodology that tracks such errors automatically and then optimizes accuracy parameters to guarantee a specified…

量子物理 · 物理学 2021-01-06 Giulia Meuli , Mathias Soeken , Martin Roetteler , Thomas Häner

Motivated by applications in service systems, we consider queueing systems where each customer must be handled by a server with the right skill set. We focus on optimizing the routing of customers to servers in order to maximize the total…

机器学习 · 计算机科学 2024-12-16 Sanne van Kempen , Jaron Sanders , Fiona Sloothaak , Maarten G. Wolf

The parameters tuning of event generators is a research topic characterized by complex choices: the generator response to parameter variations is difficult to obtain on a theoretical basis, and numerical methods are hardly tractable due to…

计算物理 · 物理学 2021-03-17 Marco Lazzarin , Simone Alioli , Stefano Carrazza

When State of Charge, State of Health, and parameters of the Lithium-ion battery are estimated simultaneously, the estimation accuracy is hard to be ensured due to uncertainties in the estimation process. To improve the estimation…

系统与控制 · 电气工程与系统科学 2024-12-20 Ziyou Song , Jun Hou , Xuefeng Li , Xiaogang Wu , Xiaosong Hu , Heath Hofmann , Jing Sun

We study both the value function and Q-function formulation of the Linear Programming approach to Approximate Dynamic Programming. The approach is model-based and optimizes over a restricted function space to approximate the value function…

系统与控制 · 计算机科学 2018-08-31 Paul N. Beuchat , Angelos Georghiou , John Lygeros

Q-learning is a regression-based approach that is widely used to formalize the development of an optimal dynamic treatment strategy. Finite dimensional working models are typically used to estimate certain nuisance parameters, and…

统计方法学 · 统计学 2020-03-30 Ashkan Ertefaie , James R. McKay , David Oslin , Robert L. Strawderman

Q-learning is a widely used reinforcement learning technique for solving path planning problems. It primarily involves the interaction between an agent and its environment, enabling the agent to learn an optimal strategy that maximizes…

机器人学 · 计算机科学 2024-12-18 Yiming Ji , Kaijie Yun , Yang Liu , Zongwu Xie , Hong Liu

In many real-world scenarios involving high-stakes and safety implications, a human decision-maker (HDM) may receive recommendations from an artificial intelligence while holding the ultimate responsibility of making decisions. In this…

机器学习 · 计算机科学 2024-07-18 Ioannis Faros , Aditya Dave , Andreas A. Malikopoulos

Model-based reinforcement learning techniques accelerate the learning task by employing a transition model to make predictions. In this paper, a model-based learning approach is presented that iteratively computes the optimal value function…

最优化与控制 · 数学 2020-10-22 Milad Farsi , Jun Liu

Hyper-parameter Tuning is among the most critical stages in building machine learning solutions. This paper demonstrates how multi-agent systems can be utilized to develop a distributed technique for determining near-optimal values for any…

机器学习 · 计算机科学 2022-05-12 Ahmad Esmaeili , Zahra Ghorrati , Eric Matson

We study a Q learning algorithm for continuous time stochastic control problems. The proposed algorithm uses the sampled state process by discretizing the state and control action spaces under piece-wise constant control processes. We show…

最优化与控制 · 数学 2023-03-10 Erhan Bayraktar , Ali Devran Kara

We consider the problem of learning control policies that optimize a reward function while satisfying constraints due to considerations of safety, fairness, or other costs. We propose a new algorithm, Projection-Based Constrained Policy…

机器学习 · 计算机科学 2020-10-08 Tsung-Yen Yang , Justinian Rosca , Karthik Narasimhan , Peter J. Ramadge

Asynchronous Q-learning aims to learn the optimal action-value function (or Q-function) of a Markov decision process (MDP), based on a single trajectory of Markovian samples induced by a behavior policy. Focusing on a $\gamma$-discounted…

机器学习 · 计算机科学 2022-09-13 Gen Li , Yuting Wei , Yuejie Chi , Yuantao Gu , Yuxin Chen
‹ 上一页 1 8 9 10 下一页 ›