中文
相关论文

相关论文: SPRIG: Stackelberg Perception-Reinforcement Learni…

200 篇论文

We propose a novel Reinforcement Learning model for discrete environments, which is inherently interpretable and supports the discovery of deep subgoal hierarchies. In the model, an agent learns information about environment in the form of…

人工智能 · 计算机科学 2022-02-16 Alexander Demin , Denis Ponomaryov

While reinforcement learning (RL) can empower autonomous agents by enabling self-improvement through interaction, its practical adoption remains challenging due to costly rollouts, limited task diversity, unreliable reward signals, and…

We study Stackelberg games where a principal repeatedly interacts with a non-myopic long-lived agent, without knowing the agent's payoff function. Although learning in Stackelberg games is well-understood when the agent is myopic, dealing…

计算机科学与博弈论 · 计算机科学 2025-05-29 Nika Haghtalab , Thodoris Lykouris , Sloan Nietert , Alexander Wei

Reinforcement Learning algorithms can learn complex behavioral patterns for sequential decision making tasks wherein an agent interacts with an environment and acquires feedback in the form of rewards sampled from it. Traditionally, such…

机器学习 · 计算机科学 2020-09-23 Sahil Sharma , Aravind Srinivas , Balaraman Ravindran

Multi-robot cooperation requires agents to make decisions that are consistent with the shared goal without disregarding action-specific preferences that might arise from asymmetry in capabilities and individual objectives. To accomplish…

机器人学 · 计算机科学 2021-05-07 Joewie J. Koh , Guohui Ding , Christoffer Heckman , Lijun Chen , Alessandro Roncone

As predictive models are deployed into the real world, they must increasingly contend with strategic behavior. A growing body of work on strategic classification treats this problem as a Stackelberg game: the decision-maker "leads" in the…

机器学习 · 计算机科学 2022-02-01 Tijana Zrnic , Eric Mazumdar , S. Shankar Sastry , Michael I. Jordan

Batch reinforcement learning (RL) defines the task of learning from a fixed batch of data lacking exhaustive exploration. Worst-case optimality algorithms, which calibrate a value-function model class from logged experience and perform some…

机器学习 · 统计学 2023-10-03 Wenzhuo Zhou , Annie Qu

We propose a new variant of the strategic classification problem: a principal reveals a classifier, and $n$ agents report their (possibly manipulated) features to be classified. Motivated by real-world applications, our model crucially…

计算机科学与博弈论 · 计算机科学 2025-02-28 Safwan Hossain , Evi Micha , Yiling Chen , Ariel Procaccia

With the rapid evolution of wireless mobile devices, there emerges an increased need to design effective collaboration mechanisms between intelligent agents, so as to gradually approach the final collective objective through continuously…

人工智能 · 计算机科学 2021-02-02 Xing Xu , Rongpeng Li , Zhifeng Zhao , Honggang Zhang

In stochastic games with incomplete information, the uncertainty is evoked by the lack of knowledge about a player's own and the other players' types, i.e. the utility function and the policy space, and also the inherent stochasticity of…

机器学习 · 计算机科学 2022-03-21 Hannes Eriksson , Debabrota Basu , Mina Alibeigi , Christos Dimitrakakis

We introduce the framework of LLM-Stackelberg games, a class of sequential decision-making models that integrate large language models (LLMs) into strategic interactions between a leader and a follower. Departing from classical Stackelberg…

人工智能 · 计算机科学 2025-07-15 Quanyan Zhu

Reinforcement Learning (RL) algorithms have been successfully applied to real world situations like illegal smuggling, poaching, deforestation, climate change, airport security, etc. These scenarios can be framed as Stackelberg security…

机器学习 · 计算机科学 2022-12-01 Saptarashmi Bandyopadhyay , Chenqi Zhu , Philip Daniel , Joshua Morrison , Ethan Shay , John Dickerson

When interacting with other decision-making agents in non-adversarial scenarios, it is critical for an autonomous agent to have inferable behavior: The agent's actions must convey their intention and strategy. We model the inferability…

计算机科学与博弈论 · 计算机科学 2025-06-03 Mustafa O. Karabag , Sophia Smith , Negar Mehr , David Fridovich-Keil , Ufuk Topcu

Learning a control policy capable of adapting to time-varying and potentially evolving system dynamics has been a great challenge to the mainstream reinforcement learning (RL). Mainly, the ever-changing system properties would continuously…

机器学习 · 计算机科学 2022-08-31 Po-Hsiang Chiu , Manfred Huber

Reinforcement learning is essential for neural architecture search and hyperparameter optimization, but the conventional approaches impede widespread use due to prohibitive time and computational costs. Inspired by DeepSeek-V3 multi-token…

机器学习 · 计算机科学 2025-06-19 Zheng Li , Jerry Cheng , Huanying Helen Gu

Machine learning systems have been widely used to make decisions about individuals who may behave strategically to receive favorable outcomes, e.g., they may genuinely improve the true labels or manipulate observable features directly to…

人工智能 · 计算机科学 2024-10-30 Tian Xie , Zhiqun Zuo , Mohammad Mahdi Khalili , Xueru Zhang

Swarm intelligence emerges from decentralised interactions among simple agents, enabling collective problem-solving. This study establishes a theoretical equivalence between pheromone-mediated aggregation in \celeg\ and reinforcement…

人工智能 · 计算机科学 2025-09-25 Aymeric Vellinger , Nemanja Antonic , Elio Tuci

Effective decision-making in partially observable environments demands robust memory management. Despite their success in supervised learning, current deep-learning memory models struggle in reinforcement learning environments that are…

机器学习 · 计算机科学 2024-10-15 Hung Le , Kien Do , Dung Nguyen , Sunil Gupta , Svetha Venkatesh

Group activities usually involve spatiotemporal dynamics among many interactive individuals, while only a few participants at several key frames essentially define the activity. Therefore, effectively modeling the group-relevant and…

计算机视觉与模式识别 · 计算机科学 2020-03-04 Guyue Hu , Bo Cui , Yuan He , Shan Yu

Deep Reinforcement Learning has enabled the control of increasingly complex and high-dimensional problems. However, the need of vast amounts of data before reasonable performance is attained prevents its widespread application. We employ…

机器学习 · 计算机科学 2020-04-08 Jan Scholten , Daan Wout , Carlos Celemin , Jens Kober