Related papers: Markovian Persuasion with Two States
We present a new behavioural distance over the state space of a Markov decision process, and demonstrate the use of this distance as an effective means of shaping the learnt representations of deep reinforcement learning agents. While…
Controllable Markov chains describe the dynamics of sequential decision making tasks and are the central component in optimal control and reinforcement learning. In this work, we give the general form of an optimal policy for learning…
The general picture of game theoretic modeling dealt with here is characterized by a set of big players, also referred to as principals or major agents, acting on the background of large pools of small players, the impact of the behavior of…
Starting from the Avellaneda-Stoikov framework, we consider a market maker who wants to optimally set bid/ask quotes over a finite time horizon, to maximize her expected utility. The intensities of the orders she receives depend not only on…
The optimal control for mobile agents is an important and challenging issue. Recent work shows that using randomized mechanism in agents' control can make the state unpredictable, and thus improve the security of agents. However, the…
This paper presents a novel state representation for reward-free Markov decision processes. The idea is to learn, in a self-supervised manner, an embedding space where distances between pairs of embedded states correspond to the minimum…
We present a model that explores the influence of persuasion in a population of agents with positive and negative opinion orientations. The opinion of each agent is represented by an integer number $k$ that expresses its level of agreement…
Analytical approaches in models of opinion formation have been extensively studied either for an opinion represented as a discrete or a continuous variable. In this paper, we analyze a model which combines both approaches. The state of an…
In this work we study the optimal execution problem with multiplicative price impact in algorithm trading, when an agent holds an initial position of shares of a financial asset. The inter-selling-decision times are modelled by the arrival…
In multi-agent reinforcement learning, the problem of learning to act is particularly difficult because the policies of co-players may be heavily conditioned on information only observed by them. On the other hand, humans readily form…
In a zero-sum stochastic game with signals, at each stage, two adversary players take decisions and receive a stage payoff determined by these decisions and a variable called state. The state follows a Markov chain, that is controlled by…
Many decision problems in economics, information technology, and industry can be transformed to an optimal stopping of adapted random vectors with some utility function over the set of Markov times with respect to filtration build by the…
We study strategic information transmission in a hierarchical setting where information gets transmitted through a chain of agents up to a decision maker whose action is of importance to every agent. This situation could arise whenever an…
General purpose intelligent learning agents cycle through (complex,non-MDP) sequences of observations, actions, and rewards. On the other hand, reinforcement learning is well-developed for small finite state Markov Decision Processes…
When humans are given a policy to execute, there can be policy execution errors and deviations in policy if there is uncertainty in identifying a state. This can happen due to the human agent's cognitive limitations and/or perceptual…
This paper looks at predictability problems, i.e., wherein an agent must choose its strategy in order to optimize the predictions that an external observer could make. We address these problems while taking into account uncertainties on the…
We study a model of binary decision making when a certain population of agents is initially seeded with two different opinions, `$+$' and `$-$', with fractions $p_1$ and $p_2$ respectively, $p_1+p_2=1$. Individuals can reverse their initial…
We develop a model-free approach to optimally control stochastic, Markovian systems subject to a reach-avoid constraint. Specifically, the state trajectory must remain within a safe set while reaching a target set within a finite time…
This paper studies a multi-player, general-sum stochastic game characterized by a dual-stage temporal structure per period. The agents face uncertainty regarding the time-evolving state that is realized at the beginning of each period.…
We present an approach to reduce the communication required between agents in a Multi-Agent learning system by exploiting the inherent robustness of the underlying Markov Decision Process. We compute so-called robustness surrogate functions…