English
Related papers

Related papers: Optimally Deceiving a Learning Leader in Stackelbe…

200 papers

We introduce the "inverse bandit" problem of estimating the rewards of a multi-armed bandit instance from observing the learning process of a low-regret demonstrator. Existing approaches to the related problem of inverse reinforcement…

Machine Learning · Statistics 2022-02-23 Wenshuo Guo , Kumar Krishna Agrawal , Aditya Grover , Vidya Muthukumar , Ashwin Pananjady

In shared autonomy, a critical tension arises when an automated assistant must choose between obeying a human's instruction and deliberately overriding it to prevent harm. This safety-critical behavior is known as intelligent disobedience.…

Artificial Intelligence · Computer Science 2026-03-24 Benedikt Hornig , Reuth Mirsky

Interdicting a criminal with limited police resources is a challenging task as the criminal changes location over time. The size of the large transportation network further adds to the difficulty of this scenario. To tackle this issue, we…

Artificial Intelligence · Computer Science 2026-04-08 Sukanya Samanta , Kei Kimura , Makoto Yokoo , Palash Dey

We consider the problem of online learning and its application to solving minimax games. For the online learning problem, Follow the Perturbed Leader (FTPL) is a widely studied algorithm which enjoys the optimal $O(T^{1/2})$ worst-case…

Machine Learning · Computer Science 2020-06-16 Arun Sai Suggala , Praneeth Netrapalli

Classical coding-theoretic guarantees often rely on trust assumptions, such as requiring sufficiently many honest nodes compared with adversarial ones. These assumptions are difficult to enforce in open decentralized systems where…

Information Theory · Computer Science 2026-05-12 Hanzaleh Akbari Nodehi , Parsa Moradi , Mohammad Ali Maddah-Ali

We consider a seller faced with buyers which have the ability to delay their decision, which we call patience. Each buyer's type is composed of value and patience, and it is sampled i.i.d. from a distribution. The seller, using posted…

Computer Science and Game Theory · Computer Science 2022-06-28 Eitan-Hai Mashiah , Idan Attias , Yishay Mansour

A ubiquitous learning problem in today's digital market is, during repeated interactions between a seller and a buyer, how a seller can gradually learn optimal pricing decisions based on the buyer's past purchase responses. A fundamental…

Computer Science and Game Theory · Computer Science 2021-10-06 Quinlan Dawkins , Minbiao Han , Haifeng Xu

This paper analyzes a class of Stackelberg games where different actors compete for shared resources and a central authority tries to balance the demand through a pricing mechanism. Situations like this can for instance occur when fleet…

Systems and Control · Electrical Eng. & Systems 2023-04-25 Marko Maljkovic , Gustav Nilsson , Nikolas Geroliminis

In this work, a novel Stackelberg game theoretic framework is proposed for trading energy bidirectionally between the demand-response (DR) aggregator and the prosumers. This formulation allows for flexible energy arbitrage and additional…

Machine Learning · Computer Science 2024-10-28 Styliani I. Kampezidou , Justin Romberg , Kyriakos G. Vamvoudakis , Dimitri N. Mavris

Demand-side management (DSM) enables distribution system operators (DSOs) to steer electricity consumption through dynamic price signals or incentive mechanisms, thereby leveraging end-users' flexibility potential for delivering grid…

Optimization and Control · Mathematics 2026-05-04 Silvia Cianchi , Reza Rahimi Baghbadorani , Anibal Sanjab , Sergio Grammatico

In this study, we explore the application of game theory, in particular Stackelberg games, to address the issue of effective coordination strategy generation for heterogeneous robots with one-way communication. To that end, focusing on the…

Robotics · Computer Science 2023-08-01 Yuhan Zhao , Baichuan Huang , Jingjin Yu , Quanyan Zhu

Reinforcement learning can greatly benefit from the use of options as a way of encoding recurring behaviours and to foster exploration. An important open problem is how can an agent autonomously learn useful options when solving particular…

Machine Learning · Computer Science 2020-01-07 Manuel Del Verme , Bruno Castro da Silva , Gianluca Baldassarre

In this technical note, we consider the linear-quadratic time-inconsistent mean-field type leader-follower Stackelberg differential game with an adapted open-loop information structure. The objective functionals of the leader and the…

Optimization and Control · Mathematics 2019-11-12 Jun Moon , Hyun Jong Yang

This paper is concerned with a Stackelberg stochastic differential game with asymmetric noisy observation, with one follower and one leader. In our model, the follower cannot observe the state process directly, but could observe a noisy…

Optimization and Control · Mathematics 2020-07-14 Yueyang Zheng , Jingtao Shi

This paper is devoted to a Stackelberg stochastic differential game for a linear mean-field type stochastic differential system with a mean-field type quadratic cost functional in finite horizon. The coefficients in the state equation and…

Optimization and Control · Mathematics 2023-08-22 Zixuan Li , Jingtao Shi

Learning algorithms are often used to make decisions in sequential decision-making environments. In multi-agent settings, the decisions of each agent can affect the utilities/losses of the other agents. Therefore, if an agent is good at…

Computer Science and Game Theory · Computer Science 2024-07-09 Angelos Assos , Yuval Dagan , Constantinos Daskalakis

Federated learning utilizes various resources provided by participants to collaboratively train a global model, which potentially address the data privacy issue of machine learning. In such promising paradigm, the performance will be…

Machine Learning · Computer Science 2021-06-30 Rongfei Zeng , Chao Zeng , Xingwei Wang , Bo Li , Xiaowen Chu

We study online learning in episodic constrained Markov decision processes (CMDPs), where the learner aims at collecting as much reward as possible over the episodes, while satisfying some long-term constraints during the learning process.…

In this work, we propose the first fully first-order method to compute an epsilon stationary Stackelberg equilibrium with convergence guarantees. To achieve this, we first reframe the leader follower interaction as single level constrained…

Optimization and Control · Mathematics 2025-09-11 April Niu , Kai Wang , Juba Ziani

Follow-The-Regularized-Leader (FTRL) is known as an effective and versatile approach in online learning, where appropriate choice of the learning rate is crucial for smaller regret. To this end, we formulate the problem of adjusting FTRL's…

Machine Learning · Computer Science 2024-03-12 Shinji Ito , Taira Tsuchiya , Junya Honda