中文
相关论文

相关论文: Reinforcement Learning for Inverse Non-Cooperative…

200 篇论文

We formulate a two-team linear quadratic stochas- tic dynamic game featuring two opposing teams each with decentralized information structures. We introduce the concept of mutual quadratic invariance (MQI), which, analogously to quadratic…

动力系统 · 数学 2016-07-20 Marcello Colombino , Roy S. Smith , Tyler H. Summers

Zero-sum games arise in a wide variety of problems, including robust optimization and adversarial learning. However, algorithms deployed for finding a local Nash equilibrium in these games often converge to non-Nash stationary points. This…

计算机科学与博弈论 · 计算机科学 2025-09-30 Kushagra Gupta , Xinjie Liu , Ross Allen , Ufuk Topcu , David Fridovich-Keil

We revisit the problem of learning in two-player zero-sum Markov games, focusing on developing an algorithm that is uncoupled, convergent, and rational, with non-asymptotic convergence rates. We start from the case of stateless matrix game…

计算机科学与博弈论 · 计算机科学 2023-11-10 Yang Cai , Haipeng Luo , Chen-Yu Wei , Weiqiang Zheng

A variety of practical problems can be modeled by the decision-making process in multi-player games where a group of self-interested players aim at optimizing their own local objectives, while the objectives depend on the actions taken by…

最优化与控制 · 数学 2023-01-09 Yuanhanqing Huang , Jianghai Hu

For a non-cooperative differential game, the value functions of the various players satisfy a system of Hamilton-Jacobi equations. In the present paper, we consider a class of infinite-horizon games with nonlinear costs exponentially…

偏微分方程分析 · 数学 2014-08-07 Alberto Bressan , Fabio S. Priuli

In recent years, stabilizing unknown dynamical systems has became a critical problem in control systems engineering. Addressing this for linear time-invariant (LTI) systems is an essential fist step towards solving similar problems for more…

最优化与控制 · 数学 2025-08-08 Xinpei Zhang , Guangyan Jia

Inverse reinforcement learning aims to infer the reward function that explains expert behavior observed through trajectories of state--action pairs. A long-standing difficulty in classical IRL is the non-uniqueness of the recovered reward:…

机器学习 · 统计学 2025-12-09 Denis Belomestny , Alexey Naumov , Sergey Samsonov

Many multi-agent interaction scenarios can be naturally modeled as noncooperative games, where each agent's decisions depend on others' future actions. However, deploying game-theoretic planners for autonomous decision-making requires a…

机器学习 · 计算机科学 2026-01-05 Yash Jain , Xinjie Liu , Lasse Peters , David Fridovich-Keil , Ufuk Topcu

This paper investigates design of noncooperative games from an optimization and control theoretic perspective. Pricing mechanisms are used as a design tool to ensure that the Nash equilibrium of a fairly general class of noncooperative…

计算机科学与博弈论 · 计算机科学 2010-07-02 Tansu Alpcan , Lacra Pavel , Nem Stefanovic

We study the convergence properties of a payoff-based higher-order version of replicator dynamics, a widely studied model in evolutionary dynamics and game-theoretic learning, in contractive games. Recent work has introduced a…

系统与控制 · 电气工程与系统科学 2026-03-20 Hassan Abdelraouf , Vijay Gupta , Jeff S. Shamma

This paper focuses on a kind of linear quadratic non-zero sum differential game driven by backward stochastic differential equation with asymmetric information, which is a natural continuation of Wang and Yu [IEEE TAC (2010) 55: 1742-1747,…

最优化与控制 · 数学 2017-03-06 Guangchen Wang , Hua Xiao , Jie Xiong

This paper considers risk-averse learning in convex games involving multiple agents that aim to minimize their individual risk of incurring significantly high costs. Specifically, the agents adopt the conditional value at risk (CVaR) as a…

最优化与控制 · 数学 2024-03-18 Zifan Wang , Yi Shen , Michael M. Zavlanos , Karl H. Johansson

This paper studies social optima and Nash games for mean field linear quadratic control systems, where subsystems are coupled via dynamics and individual costs. For the social control problem, we first obtain a set of forward-backward…

最优化与控制 · 数学 2019-04-17 Bingchang Wang , Huanshui Zhang

In this paper, the inverse reinforcement learning (IRL) problem is addressed to reconstruct the unknown cost function underlying an observed optimal policy in a model-free manner, whose online adaptation with completely off-policy system…

最优化与控制 · 数学 2025-11-20 Yibei Li , Yuexin Cao , Zhixin Liu , Lihua Xie

Reinforcement learning has been shown to be an effective strategy for automatically training policies for challenging control problems. Focusing on non-cooperative multi-agent systems, we propose a novel reinforcement learning framework for…

计算机科学与博弈论 · 计算机科学 2022-06-08 Kishor Jothimurugan , Suguman Bansal , Osbert Bastani , Rajeev Alur

In the present paper, we consider a class of two players infinite horizon differential games, with piecewise smooth costs exponentially discounted in time. Through the analysis of the value functions, we study in which cases it is possible…

偏微分方程分析 · 数学 2014-08-07 Fabio S. Priuli

A growing line of work reframes preference-based fine-tuning of large language models game-theoretically: Nash Learning from Human Feedback (NLHF) recasts the problem as a zero-sum game over policies. However, optimization is over expected…

计算机科学与博弈论 · 计算机科学 2026-05-14 Max Horwitz , Jake Gonzales , Eric Mazumdar , Lillian J. Ratliff

This paper studies two important signal processing aspects of equilibrium behavior in non-cooperative games arising in social networks, namely, reinforcement learning and detection of equilibrium play. The first part of the paper presents a…

计算机科学与博弈论 · 计算机科学 2015-01-07 Omid Namvar Gharehshiran , William Hoiles , Vikram Krishnamurthy

In common-interest stochastic games all players receive an identical payoff. Players participating in such games must learn to coordinate with each other in order to receive the highest-possible value. A number of reinforcement learning…

人工智能 · 计算机科学 2011-06-28 R. I. Brafman , M. Tennenholtz

In this paper we introduce the novel framework of distributionally robust games. These are multi-player games where each player models the state of nature using a worst-case distribution, also called adversarial distribution. Thus each…

最优化与控制 · 数学 2017-07-25 Dario Bauso , Jian Gao , Hamidou Tembine