中文
相关论文

相关论文: Trustworthiness of Optimality Condition Violation …

200 篇论文

Offline reinforcement learning (RL) aims at learning an optimal strategy using a pre-collected dataset without further interactions with the environment. While various algorithms have been proposed for offline RL in the previous literature,…

机器学习 · 计算机科学 2023-03-02 Wei Xiong , Han Zhong , Chengshuai Shi , Cong Shen , Liwei Wang , Tong Zhang

Existing studies on provably efficient algorithms for Markov games (MGs) almost exclusively build on the "optimism in the face of uncertainty" (OFU) principle. This work focuses on a different approach of posterior sampling, which is…

机器学习 · 计算机科学 2022-10-06 Wei Xiong , Han Zhong , Chengshuai Shi , Cong Shen , Tong Zhang

Partially-observable Markov decision processes (POMDPs) with discounted-sum payoff are a standard framework to model a wide range of problems related to decision making under uncertainty. Traditionally, the goal has been to obtain policies…

人工智能 · 计算机科学 2018-05-01 Krishnendu Chatterjee , Adrián Elgyütt , Petr Novotný , Owen Rouillé

We propose and study several inverse problems for the mean field games (MFG) system in a bounded domain. Our focus is on simultaneously recovering the running cost and the Hamiltonian within the MFG system by the associated boundary…

最优化与控制 · 数学 2024-03-05 Hongyu Liu , Shen Zhang

This paper presents a novel methodology to enforce motion safety guarantees even in the event of a sudden loss of control capabilities by any agent within a multi-agent system. This passive safety methodology permits the replacement of…

最优化与控制 · 数学 2023-05-29 Tommaso Guffanti , Simone D'Amico

This paper explores the use of Maximum Causal Entropy Inverse Reinforcement Learning (IRL) within the context of discrete-time stationary Mean-Field Games (MFGs) characterized by finite state spaces and an infinite-horizon,…

系统与控制 · 电气工程与系统科学 2025-07-22 Berkay Anahtarci , Can Deha Kariksiz , Naci Saldi

Pursuit-evasion scenarios appear widely in robotics, security domains, and many other real-world situations. We focus on two-player pursuit-evasion games with concurrent moves, infinite horizon, and discounted rewards. We assume that the…

计算机科学与博弈论 · 计算机科学 2016-08-05 Karel Horák , Branislav Bošanský

Regularization of control policies using entropy can be instrumental in adjusting predictability of real-world systems. Applications benefiting from such approaches range from, e.g., cybersecurity, which aims at maximal unpredictability, to…

系统与控制 · 电气工程与系统科学 2026-02-18 Menno van Zutphen , Giannis Delimpaltadakis , Maurice Heemels , Duarte Antunes

This paper investigates the optimization problem of an infinite stage discrete time Markov decision process (MDP) with a long-run average metric considering both mean and variance of rewards together. Such performance metric is important…

最优化与控制 · 数学 2020-08-11 Li Xia

Inverse Optimal Control (IOC) aims to infer the underlying cost functional of an agent from observations of its expert behavior. This paper focuses on the IOC problem within the continuous-time linear quadratic regulator framework,…

最优化与控制 · 数学 2025-07-29 Meiling Yu , Lechen Feng , Lei Jiang , Yuan-Hua Ni

We study the tracking of a trajectory for a nonholonomic system by recasting the problem as a constrained optimal control problem. The cost function is chosen to minimize the error in positions and velocities between the trajectory of a…

This paper presents a new safety specification method that is robust against errors in the probability distribution of disturbances. Our proposed distributionally robust safe policy maximizes the probability of a system remaining in a…

最优化与控制 · 数学 2018-10-05 Insoon Yang

Markov Potential Games (MPGs) form an important sub-class of Markov games, which are a common framework to model multi-agent reinforcement learning problems. In particular, MPGs include as a special case the identical-interest setting where…

机器学习 · 计算机科学 2024-08-16 Pragnya Alatur , Anas Barakat , Niao He

Decomposition, i.e. independently analyzing possible subgames, has proven to be an essential principle for effective decision-making in perfect information games. However, in imperfect information games, decomposition has proven to be…

计算机科学与博弈论 · 计算机科学 2014-04-22 Neil Burch , Michael Johanson , Michael Bowling

We study synthesis problems with constraints in partially observable Markov decision processes (POMDPs), where the objective is to compute a strategy for an agent that is guaranteed to satisfy certain safety and performance specifications.…

This paper addresses the problem of online inverse reinforcement learning for nonlinear systems with modeling uncertainties while in the presence of unknown disturbances. The developed approach observes state and input trajectories for an…

系统与控制 · 电气工程与系统科学 2021-07-07 Ryan Self , Moad Abudia , Rushikesh Kamalapurkar

In the Minority, Majority and Dollar Games (MG, MAJG, $G), synthetic agents compete for rewards, at each time-step acting in accord with the previously best-performing of their limited sets of strategies. Different components and/or aspects…

交易与市场微观结构 · 定量金融 2009-11-13 J. B. Satinover , D. Sornette

Online gradient descent (OGD) is well known to be doubly optimal under strong convexity or monotonicity assumptions: (1) in the single-agent setting, it achieves an optimal regret of $\Theta(\log T)$ for strongly convex cost functions; and…

计算机科学与博弈论 · 计算机科学 2024-04-01 Michael I. Jordan , Tianyi Lin , Zhengyuan Zhou

Game dynamics, which describe how agents' strategies evolve over time based on past interactions, can exhibit a variety of undesirable behaviours including convergence to suboptimal equilibria, cycling, and chaos. While central planners can…

系统与控制 · 电气工程与系统科学 2025-11-25 Ilayda Canyakmaz , Iosif Sakos , Wayne Lin , Antonios Varvitsiotis , Georgios Piliouras

We develop a feedback controller that minimizes the observability of a set of adversarial sensors of a linear system, while adhering to strict closed-loop performance constraints. We quantify the effectiveness of adversarial sensors using…

系统与控制 · 电气工程与系统科学 2026-02-24 Filippos Fotiadis , Ufuk Topcu