中文
相关论文

相关论文: Efficient and Safe Exploration in Deterministic Ma…

200 篇论文

This paper deals with control of partially observable discrete-time stochastic systems. It introduces and studies Markov Decision Processes with Incomplete Information and with semi-uniform Feller transition probabilities. The important…

最优化与控制 · 数学 2022-08-30 Eugene A. Feinberg , Pavlo O. Kasyanov , Michael Z. Zgurovsky

The primary goal of reinforcement learning is to develop decision-making policies that prioritize optimal performance, frequently without considering safety. In contrast, safe reinforcement learning seeks to reduce or avoid unsafe behavior.…

机器学习 · 计算机科学 2025-06-17 Zahra Shahrooei , Ali Baheri

A major challenge in deploying reinforcement learning in online tasks is ensuring that safety is maintained throughout the learning process. In this work, we propose CERL, a new method for solving constrained Markov decision processes while…

机器学习 · 计算机科学 2024-05-10 Yarden As , Bhavya Sukhija , Andreas Krause

Multi-agent autonomous exploration is essential for applications such as environmental monitoring, search and rescue, and industrial-scale surveillance. However, effective coordination under communication constraints remains a significant…

机器人学 · 计算机科学 2026-04-06 John Lewis Devassy , Meysam Basiri , Mário A. T. Figueiredo , Pedro U. Lima

Providing finite-time probabilistic safety and reach-avoid guarantees is crucial for safety-critical stochastic systems. Existing state-of-the-art barrier methods often rely on a restrictive boundedness assumption for auxiliary functions,…

系统与控制 · 电气工程与系统科学 2026-05-12 Bai Xue , Luke Ong , Dominik Wagner , Peixin Wang

Robotic manipulation in dynamic and unstructured environments requires safety mechanisms that exploit what is known and what is uncertain about the world. Existing safety filters often assume full observability, limiting their applicability…

机器人学 · 计算机科学 2025-09-17 Anna Johansson , Daniel Lindmark , Viktor Wiberg , Martin Servin

A key challenge in applying reinforcement learning to safety-critical domains is understanding how to balance exploration (needed to attain good performance on the task) with safety (needed to avoid catastrophic failure). Although a growing…

机器学习 · 计算机科学 2021-03-23 Melrose Roderick , Vaishnavh Nagarajan , J. Zico Kolter

Opacity is a generic security property, that has been defined on (non probabilistic) transition systems and later on Markov chains with labels. For a secret predicate, given as a subset of runs, and a function describing the view of an…

密码学与安全 · 计算机科学 2014-09-02 Béatrice Bérard , Krishnendu Chatterjee , Nathalie Sznajder

Addressing uncertainty is critical for autonomous systems to robustly adapt to the real world. We formulate the problem of model uncertainty as a continuous Bayes-Adaptive Markov Decision Process (BAMDP), where an agent maintains a…

机器人学 · 计算机科学 2019-05-09 Gilwoo Lee , Brian Hou , Aditya Mandalika , Jeongseok Lee , Sanjiban Choudhury , Siddhartha S. Srinivasa

Conventional navigation pipelines for legged robots remain largely geometry-centric, relying on dense SLAM representations that are fragile under rapid motion and offer limited support for semantic decision making in open-world exploration.…

机器人学 · 计算机科学 2026-03-09 Guoyang Zhao , Yudong Li , Weiqing Qi , Kai Zhang , Bonan Liu , Kai Chen , Haoang Li , Jun Ma

Partially observable Markov decision processes (POMDPs) form a prominent model for uncertainty in sequential decision making. We are interested in constructing algorithms with theoretical guarantees to determine whether the agent has a…

Autonomous systems with machine learning-based perception can exhibit unpredictable behaviors that are difficult to quantify, let alone verify. Such behaviors are convenient to capture in probabilistic models, but probabilistic model…

计算机科学中的逻辑 · 计算机科学 2022-03-17 Matthew Cleaveland , Ivan Ruchkin , Oleg Sokolsky , Insup Lee

Learning optimal control policies directly on physical systems is challenging since even a single failure can lead to costly hardware damage. Most existing model-free learning methods that guarantee safety, i.e., no failures, during…

机器学习 · 计算机科学 2023-06-13 Bhavya Sukhija , Matteo Turchetta , David Lindner , Andreas Krause , Sebastian Trimpe , Dominik Baumann

Our goal is to compute a policy that guarantees improved return over a baseline policy even when the available MDP model is inaccurate. The inaccurate model may be constructed, for example, by system identification techniques when the true…

最优化与控制 · 数学 2015-06-17 Yinlam Chow , Marek Petrik , Mohammad Ghavamzadeh

As drones and autonomous cars become more widespread it is becoming increasingly important that robots can operate safely under realistic conditions. The noisy information fed into real systems means that robots must use estimates of the…

机器人学 · 计算机科学 2017-06-01 Brian Axelrod , Leslie Pack Kaelbling , Tomás Lozano-Pérez

In this paper we solve the problem of finding a trajectory that shows that a given hybrid dynamical system with deterministic evolution leaves a given set of states considered to be safe. The algorithm combines local with global search for…

系统与控制 · 计算机科学 2014-06-25 Jan Kuřátko , Stefan Ratschan

In tabular Markov decision processes (MDPs) with perfect state observability, each trajectory provides active samples from the transition distributions conditioned on state-action pairs. Consequently, accurate model estimation depends on…

机器学习 · 计算机科学 2026-02-25 Xihe Gu , Urbashi Mitra , Tara Javidi

Planning safe trajectories under model uncertainty is a fundamental challenge. Robust planning ensures safety by considering worst-case realizations, yet ignores uncertainty reduction and leads to overly conservative behavior. Actively…

机器人学 · 计算机科学 2026-04-20 Kaleb Ben Naveed , Manveer Singh , Devansh R. Agrawal , Dimitra Panagou

The discrete class algorithm presented in this paper is an efficient simulation tool for stochastic processes governed by a reasonably small set of transition rates. The algorithm is presented, its performance compared to prevailing methods…

计算物理 · 物理学 2008-02-03 Hans E. Plesser , Dietmar Wendt

Motion planning under sensing uncertainty is critical for robots in unstructured environments to guarantee safety for both the robot and any nearby humans. Most work on planning under uncertainty does not scale to high-dimensional robots…