中文
相关论文

相关论文: Trustworthiness of Optimality Condition Violation …

200 篇论文

We build on a recently introduced geometric interpretation of Markov Decision Processes (MDPs) to analyze classical MDP-solving algorithms: Value Iteration (VI) and Policy Iteration (PI). First, we develop a geometry-based analytical…

机器学习 · 计算机科学 2025-03-07 Arsenii Mustafin , Aleksei Pakharev , Alex Olshevsky , Ioannis Ch. Paschalidis

We study the problem of learning a Nash equilibrium (NE) in an imperfect information game (IIG) through self-play. Precisely, we focus on two-player, zero-sum, episodic, tabular IIG under the perfect-recall assumption where the only…

机器学习 · 统计学 2021-06-14 Tadashi Kozuno , Pierre Ménard , Rémi Munos , Michal Valko

In this paper, a condition-based imperfect maintenance model based on piecewise deterministic Markov process (PDMP) is constructed. The degradation of the system includes two types: natural degradation and random shocks. The natural…

系统与控制 · 电气工程与系统科学 2025-03-25 Weikai Wang , Xian Chen

Min-max optimization problems involving nonconvex-nonconcave objectives have found important applications in adversarial training and other multi-agent learning settings. Yet, no known gradient descent-based method is guaranteed to converge…

机器学习 · 计算机科学 2022-10-19 Constantinos Daskalakis , Noah Golowich , Stratis Skoulakis , Manolis Zampetakis

The dynamic mode decomposition (DMD) technique extracts the dominant modes characterizing the innate dynamical behavior of the system within the measurement data. For appropriate identification of dominant modes from the measurement data,…

系统与控制 · 电气工程与系统科学 2022-10-25 G. Revati , Syed Shadab , K. Sonam , S. R. Wagh , N. M. Singh

In this letter, we study dynamic game optimal control with imperfect state observations and introduce an iterative method to find a local Nash equilibrium. The algorithm consists of an iterative procedure combining a backward recursion…

最优化与控制 · 数学 2022-06-24 Armand Jordana , Bilal Hammoud , Justin Carpentier , Ludovic Righetti

Developing and fielding complex systems requires proof that they are reliably correct with respect to their design and operating requirements. Especially for autonomous systems which exhibit unanticipated emergent behavior, fully…

软件工程 · 计算机科学 2024-02-28 Matthew Litton , Doron Drusinsky , James Bret Michael

Safety is a critical issue in learning-based robotic and autonomous systems as learned information about their environments is often unreliable and inaccurate. In this paper, we propose a risk-aware motion control tool that is robust…

机器人学 · 计算机科学 2020-03-06 Astghik Hakobyan , Insoon Yang

In this paper, we consider a modified version of the control problem in a model free Markov decision process (MDP) setting with large state and action spaces. The control problem most commonly addressed in the contemporary literature is to…

人工智能 · 计算机科学 2018-02-01 Ajin George Joseph , Shalabh Bhatnagar

Dynamic routing is one of the representative control scheme in transportation, production lines, and data transmission. In the modern context of connectivity and autonomy, routing decisions are potentially vulnerable to malicious attacks.…

系统与控制 · 电气工程与系统科学 2024-04-09 Yuzhen Zhan , Li Jin

Model checking undiscounted reachability and expected-reward properties on Markov decision processes (MDPs) is key for the verification of systems that act under uncertainty. Popular algorithms are policy iteration and variants of value…

计算机科学中的逻辑 · 计算机科学 2023-01-25 Arnd Hartmanns , Sebastian Junges , Tim Quatmann , Maximilian Weininger

Dynamic mode decomposition (DMD) is a powerful data-driven technique for construction of reduced-order models of complex dynamical systems. Multiple numerical tests have demonstrated the accuracy and efficiency of DMD, but mostly for…

数值分析 · 数学 2021-07-28 Hannah Lu , Daniel M. Tartakovsky

Reinforcement learning can acquire complex behaviors from high-level specifications. However, defining a cost function that can be optimized effectively and encodes the correct task is challenging in practice. We explore how inverse optimal…

机器学习 · 计算机科学 2016-05-30 Chelsea Finn , Sergey Levine , Pieter Abbeel

In this study, a cooperative game model is presented to schedule the day-ahead operation of multi-microgrid (MMG) systems. In the proposed model, microgrids are scheduled to achieve a global optimum for the cost of the multi-microgrid…

系统与控制 · 电气工程与系统科学 2022-01-11 Mohadese Movahednia , Hamid Karimi , Shahram Jadid

We describe a method for the identification of models for dynamical systems from observational data. The method is based on the concept of symbolic regression and uses genetic programming to evolve a system of ordinary differential…

机器学习 · 计算机科学 2021-07-14 Gabriel Kronberger , Lukas Kammerer , Michael Kommenda

State-of-the-art methods for solving 2-player zero-sum imperfect information games rely on linear programming or regret minimization, though not on dynamic programming (DP) or heuristic search (HS), while the latter are often at the core of…

人工智能 · 计算机科学 2022-10-27 Aurélien Delage , Olivier Buffet , Jilles S. Dibangoye , Abdallah Saffidine

In this paper, the inverse reinforcement learning (IRL) problem is addressed to reconstruct the unknown cost function underlying an observed optimal policy in a model-free manner, whose online adaptation with completely off-policy system…

最优化与控制 · 数学 2025-11-20 Yibei Li , Yuexin Cao , Zhixin Liu , Lihua Xie

Incomplete data are common in real-world applications. Sensors fail, records are inconsistent, and datasets collected from different sources often differ in scale, sampling rate, and quality. These differences create missing values that…

机器学习 · 计算机科学 2025-12-08 Zalish Mahmud , Anantaa Kotal , Aritran Piplai

Most of the literature on learning in games has focused on the restrictive setting where the underlying repeated game does not change over time. Much less is known about the convergence of no-regret learning algorithms in dynamic multiagent…

机器学习 · 计算机科学 2023-10-19 Ioannis Anagnostides , Ioannis Panageas , Gabriele Farina , Tuomas Sandholm

We consider nonlinear inverse problems arising in the context of parameter identification for parabolic partial differential equations (PDEs). For stable reconstructions, regularization methods such as the iteratively regularized…

‹ 上一页 1 8 9 10 下一页 ›