中文
相关论文

相关论文: Trustworthiness of Optimality Condition Violation …

200 篇论文

Modeling the purposeful behavior of imperfect agents from a small number of observations is a challenging task. When restricted to the single-agent decision-theoretic setting, inverse optimal control techniques assume that observed behavior…

计算机科学与博弈论 · 计算机科学 2013-08-19 Kevin Waugh , Brian D. Ziebart , J. Andrew Bagnell

In this paper, we consider the problem of optimizing the worst-case behavior of a partially observed system. All uncontrolled disturbances are modeled as finite-valued uncertain variables. Using the theory of cost distributions, we present…

最优化与控制 · 数学 2023-02-21 Aditya Dave , Nishanth Venkatesh , Andreas A. Malikopoulos

We consider the problem of learning to play a repeated multi-agent game with an unknown reward function. Single player online learning algorithms attain strong regret bounds when provided with full information feedback, which unfortunately…

机器学习 · 计算机科学 2019-10-29 Pier Giuseppe Sessa , Ilija Bogunovic , Maryam Kamgarpour , Andreas Krause

Autonomous agents optimize the reward function we give them. What they don't know is how hard it is for us to design a reward function that actually captures what we want. When designing the reward, we might think of some specific training…

人工智能 · 计算机科学 2020-10-08 Dylan Hadfield-Menell , Smitha Milli , Pieter Abbeel , Stuart Russell , Anca Dragan

This paper introduces Chance Constrained Gaussian Process-Motion Planning (CCGP-MP), a motion planning algorithm for robotic systems under motion and state estimate uncertainties. The paper's key idea is to capture the variations in the…

机器人学 · 计算机科学 2021-07-26 Jacob J. Johnson , Michael C. Yip

Existing game-theoretic planning methods assume that the robot knows the objective functions of the other agents a priori while, in practical scenarios, this is rarely the case. This paper introduces LUCIDGames, an inverse optimal control…

机器人学 · 计算机科学 2020-11-17 Simon Le Cleac'h , Mac Schwager , Zachary Manchester

This paper presents a Gaussian Process (GP) framework, a non-parametric technique widely acknowledged for regression and classification tasks, to address inverse problems in mean field games (MFGs). By leveraging GPs, we aim to recover…

计算机科学与博弈论 · 计算机科学 2023-12-27 Jinyan Guo , Chenchen Mou , Xianjin Yang , Chao Zhou

Off-line robot dynamic identification methods are mostly based on the use of the inverse dynamic model, which is linear with respect to the dynamic parameters. This model is sampled while the robot is tracking reference trajectories that…

机器人学 · 计算机科学 2010-09-24 Maxime Gautier , Alexandre Janot , Pierre-Olivier Vandanjon

In this paper, we address the inverse problem for linear-quadratic differential non-cooperative games with output-feedback. Given players' stabilizing feedback laws, the goal is to find cost function parameters that lead to a game for which…

最优化与控制 · 数学 2024-10-27 Emin Martirosyan , Ming Cao

We consider a Markov decision process (MDP) in which actions prescribed by the controller are executed by a separate actuator, which may behave adversarially. At each time step, the controller selects and transmits an action to the…

信息论 · 计算机科学 2025-01-29 Edoardo David Santi , Gongpu Chen , Deniz Gündüz , Asaf Cohen

Moving Target Defense (MTD) is an emerging game-changing defense strategy in cybersecurity with the goal of strengthening defenders and conversely puzzling adversaries in a network environment. The successful deployment of an MTD system can…

系统与控制 · 计算机科学 2019-05-23 Jianjun Zheng , Akbar Siami Namin

Inverse Reinforcement Learning (IRL) describes the problem of learning an unknown reward function of a Markov Decision Process (MDP) from observed behavior of an agent. Since the agent's behavior originates in its policy and MDP policies…

人工智能 · 计算机科学 2016-04-14 Michael Herman , Tobias Gindele , Jörg Wagner , Felix Schmitt , Wolfram Burgard

Dynamic Movement Primitives (DMP) are an established and efficient method for encoding robotic tasks that require adaptation based on reference motions. Typically, the nominal trajectory is obtained through Programming by Demonstration…

机器人学 · 计算机科学 2025-07-23 Giovanni Braglia , Davide Tebaldi , Luigi Biagiotti

We study how a decision-maker (DM) learns from data of unknown quality to form robust, ''general-purpose'' posterior beliefs. We develop a framework for robust learning and belief formation under a minimax-regret criterion, cast as a…

理论经济学 · 经济学 2026-02-18 Yeon-Koo Che , Longjian Li , Tianling Luo

This paper introduces two new identification methods for linear quadratic (LQ) ordinal potential differential games (OPDGs). Potential games are notable for their benefits, such as the computability and guaranteed existence of Nash…

动力系统 · 数学 2025-03-06 Balint Varga , Da Huang , Sören Hohmann

This paper studies the problem of control strategy synthesis for dynamical systems with differential constraints to fulfill a given reachability goal while satisfying a set of safety rules. Particular attention is devoted to goals that…

机器人学 · 计算机科学 2013-11-07 Luis I. Reyes Castro , Pratik Chaudhari , Jana Tumova , Sertac Karaman , Emilio Frazzoli , Daniela Rus

In this paper, we present a framework for solving continuous optimal control problems when the true system dynamics are approximated through an imperfect model. We derive a control strategy by applying Pontryagin's Minimum Principle to the…

系统与控制 · 电气工程与系统科学 2026-03-31 Panagiotis Kounatidis , Andreas A. Malikopoulos

We consider approximate dynamic programming for the infinite-horizon stationary $\gamma$-discounted optimal control problem formalized by Markov Decision Processes. While in the exact case it is known that there always exists an optimal…

最优化与控制 · 数学 2013-04-23 Boris Lesner , Bruno Scherrer

We consider Incentive Decision Processes, where a principal seeks to reduce its costs due to another agent's behavior, by offering incentives to the agent for alternate behavior. We focus on the case where a principal interacts with a…

计算机科学与博弈论 · 计算机科学 2012-10-19 Sashank J. Reddi , Emma Brunskill

We consider the problem of estimating the set of all inputs that leads a system to some particular behavior. The system is modeled by an expensive-to-evaluate function, such as a computer experiment, and we are interested in its excursion…

统计方法学 · 统计学 2021-05-10 Dario Azzimonti , David Ginsbourger , Clément Chevalier , Julien Bect , Yann Richet