中文
相关论文

相关论文: StratFormer: Adaptive Opponent Modeling and Exploi…

200 篇论文

When people play a repeated game they usually try to anticipate their opponents' moves based on past observations, and then decide what action to take next. Behavioural economics studies the mechanisms by which strategic decisions are taken…

物理与社会 · 物理学 2012-04-20 Tobias Galla

Two traditional paradigms are often used to describe the behavior of agents in multi-agent complex systems. In the first one, agents are considered to be fully rational and systems are seen as multi-player games. In the second one, agents…

计算机科学与博弈论 · 计算机科学 2016-03-17 Mickael Randour

Despite recent advances in learning-based behavioral planning for autonomous systems, decision-making in multi-task missions remains a challenging problem. For instance, a mission might require a robot to explore an unknown environment,…

机器人学 · 计算机科学 2024-12-03 Akash Karthikeyan , Yash Vardhan Pant

Offline learning has become widely used due to its ability to derive effective policies from offline datasets gathered by expert demonstrators without interacting with the environment directly. Recent research has explored various ways to…

计算机科学与博弈论 · 计算机科学 2024-03-01 Shiqi Lei , Kanghoon Lee , Linjing Li , Jinkyoo Park , Jiachen Li

Reinforcement learning (RL) has achieved remarkable success in fields like robotics and autonomous driving, but adversarial attacks designed to mislead RL systems remain challenging. Existing approaches often rely on modifying the…

机器学习 · 计算机科学 2025-07-25 Junyong Jiang , Buwei Tian , Chenxing Xu , Songze Li , Lu Dong

Offline reinforcement learning (RL) suffers from the distribution shift between the offline dataset and the online environment. In multi-agent RL (MARL), this distribution shift may arise from the nonstationary opponents in the online…

机器学习 · 计算机科学 2025-02-25 Tao Li , Juan Guevara , Xinhong Xie , Quanyan Zhu

We introduce a stochastic principal-agent model. A principal and an agent interact in a stochastic environment, each privy to observations about the state not available to the other. The principal has the power of commitment, both to elicit…

计算机科学与博弈论 · 计算机科学 2024-09-13 Jiarui Gan , Rupak Majumdar , Debmalya Mandal , Goran Radanovic

This work introduces an end-to-end graph-based agent for accelerating the computational efficiency of Benders Decomposition. The agent's policy is parameterized by a graph neural network which takes as input a bipartite graph representation…

最优化与控制 · 数学 2025-11-18 Bernard T. Agyeman , Zhe Li , Ilias Mitrai , Prodromos Daoutidis

Adversarial dynamics are a critical facet within the cyber security domain, in which there exists a co-evolution between attackers and defenders in any given threat scenario. While defenders leverage capabilities to minimize the potential…

密码学与安全 · 计算机科学 2014-08-19 Michael L. Winterrose , Kevin M. Carter

In this paper, we introduce the Behavior Structformer, a method for modeling user behavior using structured tokenization within a Transformer-based architecture. By converting tracking events into dense tokens, this approach enhances model…

计算与语言 · 计算机科学 2024-06-11 Oleg Smirnov , Labinot Polisi

Hedge has been proposed as an adaptive scheme, which guides an agent's decision in resource selection and distribution problems that can be modeled as a multi-armed bandit full information game. Such problems are encountered in the areas of…

机器学习 · 计算机科学 2018-12-10 Miltiades E. Anagnostou , Maria A. Lambrou

Learning-to-Defer (L2D) enables hybrid decision-making by routing inputs either to a predictor or to external experts. While promising, L2D is highly vulnerable to adversarial perturbations, which can not only flip predictions but also…

机器学习 · 统计学 2026-05-29 Yannis Montreuil , Letian Yu , Axel Carlier , Lai Xing Ng , Wei Tsang Ooi

The development of competitive artificial Poker playing agents has proven to be a challenge, because agents must deal with unreliable information and deception which make it essential to model the opponents in order to achieve good results.…

人工智能 · 计算机科学 2013-01-28 Luís Filipe Teófilo , Luis Paulo Reis

The problem of two companies of agents with one-step memory playing game is investigated in the context of the Iterated Prisoner's Dilemma under the partial imitation rule, where a player can imitate only those moves that he has observed in…

物理与社会 · 物理学 2011-04-01 Liangsheng Zhang , Wenjin Chen , Mathis Antony , K. Y. Szeto

Text-based game environments are challenging because agents must deal with long sequences of text, execute compositional actions using text and learn from sparse rewards. We address these challenges by proposing Language Decision…

计算与语言 · 计算机科学 2023-11-21 Nicolas Gontier , Pau Rodriguez , Issam Laradji , David Vazquez , Christopher Pal

Simultaneous reproduction of all financial stylized facts is so difficult that most existing stochastic process-based and agent-based models are unable to achieve the goal. In this study, by extending the decision-making structure of…

统计金融 · 定量金融 2019-05-22 Kei Katahira , Yu Chen , Gaku Hashimoto , Hiroshi Okuda

Dynamic representation learning plays a pivotal role in understanding the evolution of linguistic content over time. On this front both context and time dynamics as well as their interplay are of prime importance. Current approaches model…

计算与语言 · 计算机科学 2024-10-23 Talia Tseriotou , Adam Tsakalidis , Maria Liakata

Dominance is a fundamental concept in game theory. In normal-form games dominated strategies can be identified in polynomial time. As a consequence, iterative removal of dominated strategies can be performed efficiently as a preprocessing…

计算机科学与博弈论 · 计算机科学 2026-03-26 Sam Ganzfried

Meta reinforcement learning (meta RL), as a combination of meta-learning ideas and reinforcement learning (RL), enables the agent to adapt to different tasks using a few samples. However, this sampling-based adaptation also makes meta RL…

机器学习 · 计算机科学 2023-03-09 Tao Li , Haozhe Lei , Quanyan Zhu

Interactive artificial intelligence in the motion control field is an interesting topic, especially when universal knowledge is adaptive to multiple tasks and universal environments. Despite there being increasing efforts in the field of…

机器学习 · 计算机科学 2024-09-12 Luo Ji , Runji Lin