中文
相关论文

相关论文: Minimizing the Negative Side Effects of Planning w…

200 篇论文

Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intended task. We systematically study this phenomenon in coding tasks using an environment-manipulation…

机器学习 · 计算机科学 2026-04-03 Rui Wu , Ruixiang Tang

Sequential decision making using Markov Decision Process underpins many realworld applications. Both model-based and model free methods have achieved strong results in these settings. However, real-world tasks must balance reward…

机器学习 · 计算机科学 2026-04-01 Janaka Chathuranga Brahmanage , Akshat Kumar

We consider the setting of iterative learning control, or model-based policy learning in the presence of uncertain, time-varying dynamics. In this setting, we propose a new performance metric, planning regret, which replaces the standard…

机器学习 · 计算机科学 2021-03-01 Naman Agarwal , Elad Hazan , Anirudha Majumdar , Karan Singh

Sample efficiency in the face of computationally expensive simulations is a common concern in surrogate modeling. Current strategies to minimize the number of samples needed are not as effective in simulated environments with wide state…

机器学习 · 计算机科学 2025-09-03 Julen Cestero , Marco Quartulli , Marcello Restelli

Efficient planning plays a crucial role in model-based reinforcement learning. Traditionally, the main planning operation is a full backup based on the current estimates of the successor states. Consequently, its computation time is…

人工智能 · 计算机科学 2013-01-14 Harm van Seijen , Richard S. Sutton

In this paper, we consider the problem of safety assessment for Markov decision processes without explicit knowledge of the model. We aim to learn probabilistic safety specifications associated with a given policy without compromising the…

系统与控制 · 电气工程与系统科学 2023-12-11 Abhijit Mazumdar , Rafal Wisniewski , Manuela L. Bujorianu

We propose a multi-agent based computational framework for modeling decision-making and strategic interaction at micro level for smart vehicles in a smart world. The concepts of Markov game and best response dynamics are heavily leveraged.…

多智能体系统 · 计算机科学 2022-01-05 Qi Dai , Xunnong Xu , Wen Guo , Suzhou Huang , Dimitar Filev

Behavior cloning of expert demonstrations can speed up learning optimal policies in a more sample-efficient way over reinforcement learning. However, the policy cannot extrapolate well to unseen states outside of the demonstration data,…

机器学习 · 计算机科学 2022-10-19 Jung Yeon Park , Lawson L. S. Wong

A fundamental (and largely open) challenge in sequential decision-making is dealing with non-stationary environments, where exogenous environmental conditions change over time. Such problems are traditionally modeled as non-stationary…

人工智能 · 计算机科学 2024-01-23 Baiting Luo , Yunuo Zhang , Abhishek Dubey , Ayan Mukhopadhyay

In this article, we work towards the goal of developing agents that can learn to act in complex worlds. We develop a probabilistic, relational planning rule representation that compactly models noisy, nondeterministic action effects, and…

机器学习 · 计算机科学 2011-10-12 L. P. Kaelbling , H. M. Pasula , L. S. Zettlemoyer

Online planning in Markov Decision Processes (MDPs) enables agents to make sequential decisions by simulating future trajectories from the current state, making it well-suited for large-scale or dynamic environments. Sample-based methods…

人工智能 · 计算机科学 2025-09-22 Tamir Shazman , Idan Lev-Yehudi , Ron Benchetit , Vadim Indelman

We focus on the problem of long-range dynamic replanning for off-road autonomous vehicles, where a robot plans paths through a previously unobserved environment while continuously receiving noisy local observations. An effective approach…

机器人学 · 计算机科学 2024-03-19 Matt Schmittle , Rohan Baijal , Brian Hou , Siddhartha Srinivasa , Byron Boots

The possibility of flexibly assigning spectrum resources with channels of different sizes greatly improves the spectral efficiency of optical networks, but can also lead to unwanted spectrum fragmentation.We study this problem in a scenario…

网络与互联网体系结构 · 计算机科学 2018-04-30 Alexander Erreygers , Cristina Rottondi , Giacomo Verticale , Jasper De Bock

Replanning via determinization is a recent, popular approach for online planning in MDPs. In this paper we adapt this idea to classical, non-stochastic domains with partial information and sensing actions, presenting a new planner: SDR…

人工智能 · 计算机科学 2014-01-24 Ronen I. Brafman , Guy Shani

In many environments only a tiny subset of all states yield high reward. In these cases, few of the interactions with the environment provide a relevant learning signal. Hence, we may want to preferentially train on those high-reward states…

We present a state-based regression function for planning domains where an agent does not have complete information and may have sensing actions. We consider binary domains and employ a three-valued characterization of domains with sensing…

人工智能 · 计算机科学 2017-01-11 Le-Chi Tuan , Chitta Baral , Tran Cao Son

We consider a control problem for a finite-state Markov system whose performance is evaluated by a coherent Markov risk measure. For each policy, the risk of a state is approximated by a function of its features, thus leading to a…

最优化与控制 · 数学 2023-12-05 Andrzej Ruszczynski , Shangzhe Yang

In this work we address the problem of finding feasible policies for Constrained Markov Decision Processes under probability one constraints. We argue that stationary policies are not sufficient for solving this problem, and that a rich…

机器学习 · 计算机科学 2023-02-14 Agustin Castellano , Hancheng Min , Juan Bazerque , Enrique Mallada

The phenomenon of brain drain, that is the emigration of highly skilled people, has many undesirable effects, particularly for developing countries. In this study, an agent-based model is developed to understand the dynamics of such…

物理与社会 · 物理学 2021-03-30 Furkan Gürsoy , Bertan Badur

Many potential applications of reinforcement learning (RL) are stymied by the large numbers of samples required to learn an effective policy. This is especially true when applying RL to real-world control tasks, e.g. in the sciences or…