中文
相关论文

相关论文: Learning State-Dependent Policy Parametrizations f…

200 篇论文

A decision maker typically (i) incorporates training data to learn about the relative effectiveness of treatments, and (ii) chooses an implementation mechanism that implies an ``optimal'' predicted outcome distribution according to some…

计量经济学 · 经济学 2025-05-29 Anders Bredahl Kock , David Preinerstorfer

The integration of Reinforcement Learning (RL) with heuristic methods is an emerging trend for solving optimization problems, which leverages RL's ability to learn from the data generated during the search process. One promising approach is…

机器学习 · 计算机科学 2024-09-19 Arthur Müller , Lukas Vollenkemper

In order to effectively interact with or supervise a robot, humans need to have an accurate mental model of its capabilities and how it acts. Learned neural network policies make that particularly challenging. We propose an approach for…

机器人学 · 计算机科学 2018-10-19 Sandy H. Huang , Kush Bhatia , Pieter Abbeel , Anca D. Dragan

In many environmental applications, recurrent neural networks (RNNs) are often used to model physical variables with long temporal dependencies. However, due to mini-batch training, temporal relationships between training segments within…

Ride-hailing platforms have been facing the challenge of balancing demand and supply. Existing vehicle reposition techniques often treat drivers as homogeneous agents and relocate them deterministically, assuming compliance with the…

人工智能 · 计算机科学 2024-04-03 Haoyang Chen , Peiyan Sun , Qiyuan Song , Wanyuan Wang , Weiwei Wu , Wencan Zhang , Guanyu Gao , Yan Lyu

Behavioural characterizations (BCs) of decision-making agents, or their policies, are used to study outcomes of training algorithms and as part of the algorithms themselves to encourage unique policies, match expert policy or restrict…

人工智能 · 计算机科学 2021-10-29 Anssi Kanervisto , Tomi Kinnunen , Ville Hautamäki

Electric truck operations require routing decisions that remain feasible under limited battery range, long charging times, travel and energy consumption, and competition for shared charging infrastructure. These features make electric truck…

系统与控制 · 电气工程与系统科学 2026-04-30 Stavros Orfanoudakis , Ziyan Li , Ruixiao Yang , Nikolay Aristov , Pedro P. Vergara , Chuchu Fan , Elenna Dugundji

One promising approach towards effective robot decision making in complex, long-horizon tasks is to sequence together parameterized skills. We consider a setting where a robot is initially equipped with (1) a library of parameterized…

Recurrent neural networks are a widely used class of neural architectures. They have, however, two shortcomings. First, it is difficult to understand what exactly they learn. Second, they tend to work poorly on sequences requiring long-term…

机器学习 · 计算机科学 2019-05-08 Cheng Wang , Mathias Niepert

Many applied decision-making problems have a dynamic component: The policymaker needs not only to choose whom to treat, but also when to start which treatment. For example, a medical doctor may choose between postponing treatment (watchful…

统计方法学 · 统计学 2020-05-01 Xinkun Nie , Emma Brunskill , Stefan Wager

This paper addresses the problem of online learning in a dynamic setting. We consider a social network in which each individual observes a private signal about the underlying state of the world and communicates with her neighbors at each…

最优化与控制 · 数学 2013-10-02 Shahin Shahrampour , Alexander Rakhlin , Ali Jadbabaie

An important challenge in robust machine learning is when training data is provided by strategic sources who may intentionally report erroneous data for their own benefit. A line of work at the intersection of machine learning and mechanism…

计算机科学与博弈论 · 计算机科学 2024-12-24 Eric Balkanski , Cherlin Zhu

Parameter regularization or allocation methods are effective in overcoming catastrophic forgetting in lifelong learning. However, they solve all tasks in a sequence uniformly and ignore the differences in the learning difficulty of…

机器学习 · 计算机科学 2023-04-12 Wenjin Wang , Yunqing Hu , Qianglong Chen , Yin Zhang

Parameter control and dynamic algorithm configuration study how to dynamically choose suitable configurations of a parametrized algorithm during the optimization process. Despite being an intensively researched topic in evolutionary…

神经与进化计算 · 计算机科学 2025-07-14 Gianluca Covini , Denis Antipov , Carola Doerr

In most practical applications of reinforcement learning, it is untenable to maintain direct estimates for individual states; in continuous-state systems, it is impossible. Instead, researchers often leverage state similarity (whether…

机器学习 · 计算机科学 2021-02-03 Charline Le Lan , Marc G. Bellemare , Pablo Samuel Castro

Deep reinforcement learning (DRL) has been used to learn effective heuristics for solving complex combinatorial optimisation problem via policy networks and have demonstrated promising performance. Existing works have focused on solving…

机器学习 · 计算机科学 2020-12-25 Nasrin Sultana , Jeffrey Chan , A. K. Qin , Tabinda Sarwar

Linear dynamical systems that obey stochastic differential equations are canonical models. While optimal control of known systems has a rich literature, the problem is technically hard under model uncertainty and there are hardly any…

系统与控制 · 电气工程与系统科学 2023-06-09 Mohamad Kazem Shirani Faradonbeh , Mohamad Sadegh Shirani Faradonbeh

Autonomous agents often require multiple strategies to solve complex tasks, but determining when to switch between strategies remains challenging. This research introduces a reinforcement learning technique to learn switching thresholds…

机器学习 · 计算机科学 2025-12-09 Chris Tava

Learning per-domain generalizing policies is a key challenge in learning for planning. Standard approaches learn state-value functions represented as graph neural networks using supervised learning on optimal plans generated by a teacher…

人工智能 · 计算机科学 2026-03-19 Nicola J. Müller , Moritz Oster , Isabel Valera , Jörg Hoffmann , Timo P. Gros

In safe Reinforcement Learning (RL), safety cost is typically defined as a function dependent on the immediate state and actions. In practice, safety constraints can often be non-Markovian due to the insufficient fidelity of state…

机器学习 · 计算机科学 2024-05-07 Siow Meng Low , Akshat Kumar