中文
相关论文

相关论文: Iterative Best Response for Multi-Body Asset-Guard…

200 篇论文

Combinatorial bandits with semi-bandit feedback generalize multi-armed bandits, where the agent chooses sets of arms and observes a noisy reward for each arm contained in the chosen set. The action set satisfies a given structure such as…

机器学习 · 统计学 2021-01-22 Marc Jourdan , Mojmír Mutný , Johannes Kirschner , Andreas Krause

In this paper, we address a pursuit-evasion game involving multiple players by utilizing tools and techniques from reinforcement learning and matrix game theory. In particular, we consider the problem of steering an evader to a goal…

系统与控制 · 电气工程与系统科学 2020-03-10 Jhanani Selvakumar , Efstathios Bakolas

Many real-world systems often involve physical components or operating environments with highly nonlinear and uncertain dynamics. A number of different control algorithms can be used to design optimal controllers for such systems, assuming…

系统与控制 · 电气工程与系统科学 2023-04-06 Navid Hashemi , Justin Ruths , Jyotirmoy V. Deshmukh

This paper offers a unified perspective on different approaches to the solution of optimal control problems through the lens of constrained sequential quadratic programming. In particular, it allows us to find the relationships between…

最优化与控制 · 数学 2025-10-07 Abhijeet , Suman Chakravorty

Dynamic games are an effective paradigm for dealing with the control of multiple interacting actors. This paper introduces ALGAMES (Augmented Lagrangian GAME-theoretic Solver), a solver that handles trajectory optimization problems with…

机器人学 · 计算机科学 2021-06-01 Simon Le Cleac'h , Mac Schwager , Zachary Manchester

This paper introduces two metrics (cycle-based and memory-based metrics), grounded on a dynamical game-theoretic solution concept called sink equilibrium, for the evaluation, ranking, and computation of policies in multi-agent learning. We…

计算机科学与博弈论 · 计算机科学 2020-06-23 Rui Yan , Xiaoming Duan , Zongying Shi , Yisheng Zhong , Jason R. Marden , Francesco Bullo

We consider a sequential inspection game where an inspector uses a limited number of inspections over a larger number of time periods to detect a violation (an illegal act) of an inspectee. Compared with earlier models, we allow varying…

计算机科学与博弈论 · 计算机科学 2016-08-24 Bernhard von Stengel

We present successive convexification, a real-time-capable solution method for nonconvex trajectory optimization, with continuous-time constraint satisfaction and guaranteed convergence, that only requires first-order information. The…

最优化与控制 · 数学 2024-04-26 Purnanand Elango , Dayou Luo , Abhinav G. Kamath , Samet Uzun , Taewan Kim , Behçet Açıkmeşe

We propose a new approach for solving combinatorial optimization problem by utilizing the mechanism of chases and escapes, which has a long history in mathematics. In addition to the well-used steepest descent and neighboring search, we…

人工智能 · 计算机科学 2018-04-25 Toru Ohira

We present a provably optimal differentially private algorithm for the stochastic multi-arm bandit problem, as opposed to the private analogue of the UCB-algorithm [Mishra and Thakurta, 2015; Tossou and Dimitrakakis, 2016] which doesn't…

机器学习 · 统计学 2019-05-24 Touqir Sajed , Or Sheffet

We develop an optimization-based framework for joint real-time trajectory planning and feedback control of feedback-linearizable systems. To achieve this goal, we define a target trajectory as the optimal solution of a time-varying…

系统与控制 · 电气工程与系统科学 2020-03-17 Tianqi Zheng , John Simpson-Porco , Enrique Mallada

We present a robust framework with computational algorithms to support decision makers in sequential games. Our framework includes methods to solve games with complete information, assess the robustness of such solutions and, finally,…

统计计算 · 统计学 2024-02-22 Tahir Ekin , Roi Naveiro , Alberto Torres-Barrán , David Ríos-Insua

Recent advances in bandit tools and techniques for sequential learning are steadily enabling new applications and are promising the resolution of a range of challenging related problems. We study the game tree search problem, where the goal…

机器学习 · 统计学 2017-11-07 Emilie Kaufmann , Wouter Koolen

In this paper we investigate a differential game in which countably many dynamical objects pursue a single one. All the players perform simple motions. The duration of the game is fixed. The controls of a group of pursuers are subject to…

最优化与控制 · 数学 2014-10-10 Mehdi Salimi , Gafurjan Ibragimov , Stefan Siegmund , Somayeh Sharifi

We extend the adversarial/non-stochastic multi-play multi-armed bandit (MPMAB) to the case where the number of arms to play is variable. The work is motivated by the fact that the resources allocated to scan different critical locations in…

机器学习 · 计算机科学 2021-10-28 Yiyang Wang , Neda Masoud

This paper presents scalable algorithms for computing pure Nash equilibria (PNEs) in large-scale integer programming games (IPGs), where existing exact methods typically handle only small numbers of players. Motivated by a county-level…

计算机科学与博弈论 · 计算机科学 2026-02-26 Hyunwoo Lee , Robert Hildebrand , Wenbo Cai , İ. Esra Büyüktahtakın

The increasing prevalence of security attacks on software-intensive systems calls for new, effective methods for detecting and responding to these attacks. As one promising approach, game theory provides analytical tools for modeling the…

软件工程 · 计算机科学 2021-12-15 Mingyue Zhang , Nianyu Li , Sridhar Adepu , Eunsuk Kang , Zhi Jin

In this work we develop a numerical method for solving a type of convex graph-structured tensor optimization problems. This type of problems, which can be seen as a generalization of multi-marginal optimal transport problems with…

最优化与控制 · 数学 2024-03-25 Axel Ringh , Isabel Haasler , Yongxin Chen , Johan Karlsson

Dynamic zero-sum games are an important class of problems with applications ranging from evasion-pursuit and heads-up poker to certain adversarial versions of control problems such as multi-armed bandit and multiclass queuing problems.…

计算机科学与博弈论 · 计算机科学 2015-06-12 Martin Haugh , Chun Wang

We study a simple motion differential game of many pursuers and one evader in the plane. We give a nonempty closed convex set in the plane, and the pursuers and evader move on this set. They cannot leave this set during the game. Control…

最优化与控制 · 数学 2015-05-04 Idham Arif Alias , Gafurjan Ibragimov , Massimiliano Ferrara , Mehdi Salimi , Mansor Monsi