中文
相关论文

相关论文: Multi-Head Attention Is a Multi-Player Game

200 篇论文

Mirror play (MP) is a well-accepted primal-dual multi-agent learning algorithm where all agents simultaneously implement mirror descent in a distributed fashion. The advantage of MP over vanilla gradient play lies in its usage of mirror…

计算机科学与博弈论 · 计算机科学 2024-03-26 Yunian Pan , Tao Li , Quanyan Zhu

Recent advances in Artificial Intelligence have produced agents that can beat human world champions at games like Go, Starcraft, and Dota2. However, most of these models do not seem to play in a human-like manner: People infer others'…

人工智能 · 计算机科学 2020-08-03 Terence X. Lim , Sidney Tio , Desmond C. Ong

In addressing the challenge of exponential scaling with the number of agents we adopt a cluster-based representation to approximately solve asymmetric games of very many players. A cluster groups together agents with a similar "strategic…

计算机科学与博弈论 · 计算机科学 2012-06-18 Sevan G. Ficici , David C. Parkes , Avi Pfeffer

Recent work on Transformer-based large language models (LLMs) has revealed striking limits in their working memory capacity, similar to what has been found in human behavioral studies. Specifically, these models' performance drops…

计算与语言 · 计算机科学 2024-11-19 Dongyu Gong , Hantao Zhang

We consider a slotted-ALOHA LAN with loss-averse, noncooperative greedy users. To avoid non-Pareto equilibria, particularly deadlock, we assume probabilistic loss-averse behavior. This behavior is modeled as a modulated white noise term, in…

计算机科学与博弈论 · 计算机科学 2012-02-28 George Kesidis , Youngmi Jin

There is growing experimental evidence that $Q$-learning agents may learn to charge supracompetitive prices. We provide the first theoretical explanation for this behavior in infinite repeated games. Firms update their pricing policies…

综合经济学 · 经济学 2025-05-30 Cristian Chica , Yinglong Guo , Gilad Lerman

OpenAI has recently argued that hallucinations in large language models result primarily from misaligned evaluation incentives that reward confident guessing rather than epistemic humility. On this view, hallucination is a contingent…

计算与语言 · 计算机科学 2025-12-18 Richard Ackermann , Simeon Emanuilov

Hallucination continues to pose a major obstacle in the reasoning capabilities of large language models (LLMs). Although the Multi-Agent Debate (MAD) paradigm offers a promising solution by promoting consensus among multiple agents to…

人工智能 · 计算机科学 2025-11-17 Dayong Liang , Xiao-Yong Wei , Changmeng Zheng

Uncertainty estimation is a necessary component when implementing AI in high-risk settings, such as autonomous cars, medicine, or insurances. Large Language Models (LLMs) have seen a surge in popularity in recent years, but they are subject…

机器学习 · 计算机科学 2024-12-09 Gabriel Y. Arteaga , Thomas B. Schön , Nicolas Pielawski

We study a two-player Stackelberg game with incomplete information such that the follower's strategy belongs to a known family of parameterized functions with an unknown parameter vector. We design an adaptive learning approach to…

计算机科学与博弈论 · 计算机科学 2021-01-12 Guosong Yang , Radha Poovendran , João P. Hespanha

We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework consists of five *structurally separate* utility heads --…

人工智能 · 计算机科学 2025-11-20 Aran Nayebi

This paper proposes a new lens for studying threshold games played on networks when the thresholds are heterogeneous. These are games where agents have two possible actions, and prefer action 1 if and only if enough of their neighbours…

理论经济学 · 经济学 2025-08-07 Alastair Langtry , Sarah Taylor , Yifan Zhang

In this paper, we formulate an evolutionarymultiple access control game with continuousvariable actions and coupled constraints. We characterize equilibria of the game and show that the pure equilibria are Pareto optimal and also resilient…

计算机科学与博弈论 · 计算机科学 2015-03-19 Quanyan Zhu , Hamidou Tembine , Tamer Basar

Recent progress in Large Language Model (LLM) reasoning is increasingly driven by the refinement of post-training loss functions and alignment strategies. However, standard Reinforcement Learning (RL) paradigms like Group Relative Policy…

机器学习 · 计算机科学 2026-01-28 Kishan Panaganti , Zhenwen Liang , Wenhao Yu , Haitao Mi , Dong Yu

Game theory has been developed by scientists as a theory of strategic interaction among players who are supposed to be perfectly rational. These strategic interactions might have been presented in an auction, a business negotiation, a chess…

计算机科学与博弈论 · 计算机科学 2020-04-07 Medet Kanmaz , Elif Surer

We consider the capacitated selfish replication (CSR) game with binary preferences, over general undirected networks. We first show that such games have an associated ordinary potential function, and hence always admit a pure-strategy Nash…

计算机科学与博弈论 · 计算机科学 2016-03-14 Seyed Rasoul Etesami , Tamer Basar

In structured decision-making workflows such as form filling, compliance checking, and maintenance reporting, LLM outputs must be locally correct, globally consistent, and auditable against task-specific rules. Existing refinement methods…

机器学习 · 计算机科学 2026-05-12 Fei Xu Yu , Zuyuan Zhang , Mahdi Imani , Nathaniel D. Bastian , Tian Lan

We consider the problem of designing network cost-sharing protocols with good equilibria under uncertainty. The underlying game is a multicast game in a rooted undirected graph with nonnegative edge costs. A set of k terminal vertices or…

计算机科学与博弈论 · 计算机科学 2015-07-27 George Christodoulou , Alkmini Sgouritsa

Many real-world multi-agent interactions consider multiple distinct criteria, i.e. the payoffs are multi-objective in nature. However, the same multi-objective payoff vector may lead to different utilities for each participant. Therefore,…

多智能体系统 · 计算机科学 2020-11-17 Roxana Rădulescu , Timothy Verstraeten , Yijie Zhang , Patrick Mannion , Diederik M. Roijers , Ann Nowé

Consider a 2-player normal-form game repeated over time. We introduce an adaptive learning procedure, where the players only observe their own realized payoff at each stage. We assume that agents do not know their own payoff function, and…

计算机科学与博弈论 · 计算机科学 2013-06-13 Mario Bravo , Mathieu Faure
‹ 上一页 1 8 9 10 下一页 ›