中文
相关论文

相关论文: Stackelberg Actor-Critic: Game-Theoretic Reinforce…

200 篇论文

Traditional centralized multi-agent reinforcement learning (MARL) algorithms are sometimes unpractical in complicated applications, due to non-interactivity between agents, curse of dimensionality and computation complexity. Hence, several…

机器学习 · 计算机科学 2023-07-10 Wenhao Li , Bo Jin , Xiangfeng Wang , Junchi Yan , Hongyuan Zha

Large language model (LLM) agents have shown remarkable progress in social deduction games (SDGs). However, existing approaches primarily focus on information processing and strategy selection, overlooking the significance of persuasive…

人工智能 · 计算机科学 2026-04-15 Zhang Zheng , Deheng Ye , Peilin Zhao , Hao Wang

Actor-critic algorithms are widely used in reinforcement learning, but are challenging to mathematically analyse due to the online arrival of non-i.i.d. data samples. The distribution of the data samples dynamically changes as the model is…

机器学习 · 计算机科学 2023-09-20 Ziheng Wang , Justin Sirignano

Extracting relevant information from a stream of high-dimensional observations is a central challenge for deep reinforcement learning agents. Actor-critic algorithms add further complexity to this challenge, as it is often unclear whether…

We establish the convergence of the deep actor-critic reinforcement learning algorithm presented in [Angiuli et al., 2023a] in the setting of continuous state and action spaces with an infinite discrete-time horizon. This algorithm provides…

最优化与控制 · 数学 2025-11-11 Jean-Pierre Fouque , Mathieu Laurière , Mengrui Zhang

Information uncertainty is one of the major challenges facing applications of game theory. In the context of Stackelberg games, various approaches have been proposed to deal with the leader's incomplete knowledge about the follower's…

计算机科学与博弈论 · 计算机科学 2019-05-21 Jiarui Gan , Haifeng Xu , Qingyu Guo , Long Tran-Thanh , Zinovi Rabinovich , Michael Wooldridge

In this paper, we consider a discrete-time stochastic Stackelberg game with a single leader and multiple followers. Both the followers and the leader together have conditionally independent private types, conditioned on action and previous…

最优化与控制 · 数学 2022-09-21 Deepanshu Vasal

This paper extends off-policy reinforcement learning to the multi-agent case in which a set of networked agents communicating with their neighbors according to a time-varying graph collaboratively evaluates and improves a target policy…

机器学习 · 计算机科学 2019-11-20 Wesley Suttle , Zhuoran Yang , Kaiqing Zhang , Zhaoran Wang , Tamer Basar , Ji Liu

While deep reinforcement learning has achieved tremendous successes in various applications, most existing works only focus on maximizing the expected value of total return and thus ignore its inherent stochasticity. Such stochasticity is…

机器学习 · 计算机科学 2023-09-19 Han Zhong , Xun Deng , Ethan X. Fang , Zhuoran Yang , Zhaoran Wang , Runze Li

In this paper, we are concerned with the stabilizatbility of Stackelberg game-based systems. In particular, two players are involved in the system where one is the follower to minimize the related cost function and the other is the leader…

最优化与控制 · 数学 2021-05-04 Yue Sun , Juanjuan Xu , Huanshui Zhang

Many popular practical reinforcement learning (RL) algorithms employ evolving reward functions-through techniques such as reward shaping, entropy regularization, or curriculum learning-yet their theoretical foundations remain…

机器学习 · 计算机科学 2025-10-15 Rui Hu , Yu Chen , Longbo Huang

Deep reinforcement learning agents often face challenges to effectively coordinate perception and decision-making components, particularly in environments with high-dimensional sensory inputs where feature relevance varies. This work…

人工智能 · 计算机科学 2025-02-21 Fernando Martinez-Lopez , Juntao Chen , Yingdong Lu

This paper introduces a novel framework for modeling interacting humans in a multi-stage game. This "iterated semi network-form game" framework has the following desirable characteristics: (1) Bounded rational players, (2) strategic players…

多智能体系统 · 计算机科学 2012-07-05 Ritchie Lee , David H. Wolpert , James Bono , Scott Backhaus , Russell Bent , Brendan Tracey

University engineering capstone projects involve sustained interaction among students, faculty, and industry sponsors whose objectives are only partially aligned. While capstones are widely used in engineering education, existing analyses…

计算机与社会 · 计算机科学 2026-01-16 Richard Q. Blackwell , Eman Hammad , Congrui Jin , Jisoo Park , Albert E. Patterson

A Stackelberg game is played between a leader and a follower. The leader first chooses an action, then the follower plays his best response. The goal of the leader is to pick the action that will maximize his payoff given the follower's…

数据结构与算法 · 计算机科学 2015-11-19 Aaron Roth , Jonathan Ullman , Zhiwei Steven Wu

In this paper, the two-player leader-follower game with private inputs for feedback Stackelberg strategy is considered. In particular, the follower shares its measurement information with the leader except its historical control inputs…

最优化与控制 · 数学 2023-09-18 Yue Sun , Hongdan Li , Huanshui Zhang

We provide a general approach to reformulating any continuous-time stochastic Stackelberg differential game under closed-loop strategies as a single-level optimisation problem with target constraints. More precisely, we consider a…

最优化与控制 · 数学 2026-05-14 Camilo Hernández , Nicolás Hernández Santibáñez , Emma Hubert , Dylan Possamaï

We propose projection-free sequential algorithms for linear-quadratic dynamics games. These policy gradient based algorithms are akin to Stackelberg leadership model and can be extended to model-free settings. We show that if the leader…

系统与控制 · 电气工程与系统科学 2019-11-13 Jingjing Bu , Lillian J. Ratliff , Mehran Mesbahi

We study a new two-time-scale stochastic gradient method for solving optimization problems, where the gradients are computed with the aid of an auxiliary variable under samples generated by time-varying MDPs controlled by the underlying…

最优化与控制 · 数学 2024-08-27 Sihan Zeng , Thinh T. Doan , Justin Romberg

Solving feedback Stackelberg games with nonlinear dynamics and coupled constraints, a common scenario in practice, presents significant challenges. This work introduces an efficient method for computing approximate local feedback…

最优化与控制 · 数学 2025-04-03 Jingqi Li , Somayeh Sojoudi , Claire Tomlin , David Fridovich-Keil