中文
相关论文

相关论文: Learning to Cooperate via Policy Search

200 篇论文

Prediction is a well-studied machine learning task, and prediction algorithms are core ingredients in online products and services. Despite their centrality in the competition between online companies who offer prediction-based products,…

计算机科学与博弈论 · 计算机科学 2019-05-09 Omer Ben-Porat , Moshe Tennenholtz

Concurrent stochastic games (CSGs) are an ideal formalism for modelling probabilistic systems that feature multiple players or components with distinct objectives making concurrent, rational decisions. Examples include communication or…

计算机科学中的逻辑 · 计算机科学 2020-07-27 Marta Kwiatkowska , Gethin Norman , David Parker , Gabriel Santos

Reinforcement learning and Imitation Learning approaches utilize policy learning strategies that are difficult to generalize well with just a few examples of a task. In this work, we propose a language-conditioned semantic search-based…

机器人学 · 计算机科学 2023-12-12 Jannik Sheikh , Andrew Melnik , Gora Chand Nandi , Robert Haschke

The standard Reinforcement Learning from Human Feedback (RLHF) framework primarily focuses on optimizing the performance of large language models using pre-collected prompts. However, collecting prompts that provide comprehensive coverage…

The framework of multi-agent learning explores the dynamics of how individual agent strategies evolve in response to the evolving strategies of other agents. Of particular interest is whether or not agent strategies converge to well known…

计算机科学与博弈论 · 计算机科学 2023-11-21 Sarah A. Toonsi , Jeff S. Shamma

While it is known that shared quantum entanglement can offer improved solutions to a number of purely cooperative tasks for groups of remote agents, controversy remains regarding the legitimacy of quantum games in a competitive setting--in…

量子物理 · 物理学 2015-05-14 Charles D. Hill , Adrian P. Flitney , Nicolas C. Menicucci

Exploration is critical to a reinforcement learning agent's performance in its given environment. Prior exploration methods are often based on using heuristic auxiliary predictions to guide policy behavior, lacking a mathematically-grounded…

机器学习 · 计算机科学 2020-03-02 Lisa Lee , Benjamin Eysenbach , Emilio Parisotto , Eric Xing , Sergey Levine , Ruslan Salakhutdinov

In this work, we study the system of interacting non-cooperative two Q-learning agents, where one agent has the privilege of observing the other's actions. We show that this information asymmetry can lead to a stable outcome of population…

机器学习 · 计算机科学 2021-01-26 Ezra Tampubolon , Haris Ceribasic , Holger Boche

Probabilistic model checking for stochastic games enables formal verification of systems that comprise competing or collaborating entities operating in a stochastic environment. Despite good progress in the area, existing approaches focus…

计算机科学中的逻辑 · 计算机科学 2019-07-09 Marta Kwiatkowska , Gethin Norman , David Parker , Gabriel Santos

Reinforcement learning in partially observable domains is challenging due to the lack of observable state information. Thankfully, learning offline in a simulator with such state information is often possible. In particular, we propose a…

机器人学 · 计算机科学 2022-11-11 Hai Nguyen , Andrea Baisero , Dian Wang , Christopher Amato , Robert Platt

This paper investigates the distributed Nash equilibrium seeking problem for two-network zero-sum games with set constraints, where the two networks have the opposite nonsmooth cost functions. The interaction of the agents in each network…

最优化与控制 · 数学 2019-12-03 Dandan Yue , Ziyang Meng

Distributed Support Vector Machines (DSVM) have been developed to solve large-scale classification problems in networked systems with a large number of sensors and control units. However, the systems become more vulnerable as detection and…

机器学习 · 统计学 2018-03-14 Rui Zhang , Quanyan Zhu

Multi-agent reinforcement learning in mixed-motive settings presents a fundamental challenge: agents must balance individual interests with collective goals, which are neither fully aligned nor strictly opposed. To address this, reward…

多智能体系统 · 计算机科学 2025-08-26 Woojun Kim , Katia Sycara

In this paper, we investigate the seeking of Nash equilibrium (NE) in a non-cooperative quadratic game where all agents exchange their delayed strategy information with their neighbors. To extend best-response algorithms to the delayed…

系统与控制 · 电气工程与系统科学 2026-02-24 Kaichen Jiang , Yuyue Yan , Mingda Yue , Yuhu Wu

We consider the problem of computing mixed Nash equilibria of two-player zero-sum games with continuous sets of pure strategies and with first-order access to the payoff function. This problem arises for example in game-theory-inspired…

最优化与控制 · 数学 2025-09-04 Guillaume Wang , Lénaïc Chizat

Surveillance-Evasion (SE) games form an important class of adversarial trajectory-planning problems. We consider time-dependent SE games, in which an Evader is trying to reach its target while minimizing the cumulative exposure to a moving…

最优化与控制 · 数学 2019-09-09 Elliot Cartee , Lexiao Lai , Qianli Song , Alexander Vladimirsky

The behaviour of multi-agent learning in competitive settings is often considered under the restrictive assumption of a zero-sum game. Only under this strict requirement is the behaviour of learning well understood; beyond this, learning…

计算机科学与博弈论 · 计算机科学 2023-07-27 Aamal Hussain , Francesco Belardinelli , Georgios Piliouras

Addressing the question of how to achieve optimal decision-making under risk and uncertainty is crucial for enhancing the capabilities of artificial agents that collaborate with or support humans. In this work, we address this question in…

多智能体系统 · 计算机科学 2024-08-02 Nicole Orzan , Erman Acar , Davide Grossi , Patrick Mannion , Roxana Rădulescu

This paper considers convex games involving multiple agents that aim to minimize their own cost functions using locally available information. A common assumption in the study of such games is that the agents are symmetric, meaning that…

最优化与控制 · 数学 2025-09-25 Zifan Wang , Xinlei Yi , Yi Shen , Michael M. Zavlanos , Karl H. Johansson

Consider an important meeting to be held in a team-based organization. Taking availability constraints into account, an online scheduling poll is being used in order to decide upon the exact time of the meeting. Decisions are to be taken…

计算机科学与博弈论 · 计算机科学 2016-11-29 Robert Bredereck , Jiehua Chen , Rolf Niedermeier , Svetlana Obraztsova , Nimrod Talmon
‹ 上一页 1 8 9 10 下一页 ›