Related papers: Persuasion and Optimal Stopping
We consider the problem of stopping a diffusion process with a payoff functional that renders the problem time-inconsistent. We study stopping decisions of naive agents who reoptimize continuously in time, as well as equilibrium strategies…
This paper concerns sequential hypothesis testing in competitive multi-agent systems where agents exchange potentially manipulated information. Specifically, a two-agent scenario is studied where each agent aims to correctly infer the true…
This paper explores continuous-time and state-space optimal stopping problems from a reinforcement learning perspective. We begin by formulating the stopping problem using randomized stopping times, where the decision maker's control is…
We address the fundamental problem of selection under uncertainty by modeling it from the perspective of Bayesian persuasion. In our model, a decision maker with imperfect information always selects the option with the highest expected…
In this work an opinion formation model with heterogeneous agents is proposed. Each agent is supposed to have different power of persuasion, and besides its own level of zealotry, that is, an individual willingness to being convinced by…
A designer relies on an experimenter to provide information to a decision maker, but the experimenter has incentives to persuade rather than merely transmit information. Anticipating this motive, the designer can restrict the set of…
We introduce and study the problem of detecting whether an agent is updating their prior beliefs given new evidence in an optimal way that is Bayesian, or whether they are biased towards their own prior. In our model, biased agents form…
In many practical problems, a learning agent may want to learn the best action in hindsight without ever taking a bad action, which is significantly worse than the default production action. In general, this is impossible because the agent…
The primary goal of reinforcement learning is to develop decision-making policies that prioritize optimal performance, frequently without considering safety. In contrast, safe reinforcement learning seeks to reduce or avoid unsafe behavior.…
Motivating careerists is challenging for political organizations. Without explicit contracts, careerists often pander to public opinions or their superiors' preferences. Worse, when tasked with implementing these distorted decisions, they…
We investigate the task of learning to follow natural language instructions by jointly reasoning with visual observations and language inputs. In contrast to existing methods which start with learning from demonstrations (LfD) and then use…
We study how a principal should optimally choose between implementing a new policy and maintaining the status quo when information relevant for the decision is privately held by agents. Agents are strategic in revealing their information;…
We introduce \emph{informational punishment} to the design of mechanisms that compete with an exogenous status quo mechanism: Players can send garbled public messages with some delay, and others cannot commit to ignoring them. Optimal…
In a multi-agent system, an agent's optimal policy will typically depend on the policies chosen by others. Therefore, a key issue in multi-agent systems research is that of predicting the behaviours of others, and responding promptly to…
In this paper we investigate the potential for persuasion arising from the quantum indeterminacy of a decision-maker's beliefs, a feature that has been proposed as a formal expression of well-known cognitive limitations. We focus on a…
What is a good exploration strategy for an agent that interacts with an environment in the absence of external rewards? Ideally, we would like to get a policy driving towards a uniform state-action visitation (highly exploring) in a minimum…
Sellers in online markets face the challenge of determining the right time to sell in view of uncertain future offers. Classical stopping theory assumes that sellers have full knowledge of the value distributions, and leverage this…
A principal and $n\ge 2$ agents can launch a project if the principal proposes it and at least $k$ agents accept. Their individual payoffs from the project depend on an ex ante unknown state. The principal can conduct a test to learn about…
We study a setting in which a principal selects an agent to execute a collection of tasks according to a specified priority sequence. Agents, however, have their own individual priority sequences according to which they wish to execute the…
A temporally abstract action, or an option, is specified by a policy and a termination condition: the policy guides option behavior, and the termination condition roughly determines its length. Generally, learning with longer options (like…