Related papers: Selfish optimization and collective learning in po…
We study learning by privately informed forward-looking agents in a simple repeated-action setting of social learning. Under a symmetric signal structure, forward-looking agents behave myopically for any degrees of patience. Myopic…
We consider the problem of learning to exploit learning algorithms through repeated interactions in games. Specifically, we focus on the case of repeated two player, finite-action games, in which an optimizer aims to steer a no-regret…
We consider multi-armed bandit problems in social groups wherein each individual has bounded memory and shares the common goal of learning the best arm/option. We say an individual learns the best option if eventually (as $t \to \infty$) it…
Self-play reinforcement learning has achieved state-of-the-art, and often superhuman, performance in a variety of zero-sum games. Yet prior work has found that policies that are highly capable against regular opponents can fail…
Machine learning algorithms have reached mainstream status and are widely deployed in many applications. The accuracy of such algorithms depends significantly on the size of the underlying training dataset; in reality a small or medium…
Social dilemmas are situations in which collective welfare is at odds with individual gain. One widely studied example, due to the conflict it poses between human behaviour and game theoretic reasoning, is the Traveler's Dilemma. The…
Here we study the effects of adopting different strategies against different opponent instead of adopting the same strategy against all of them in the prisoner dilemma structured in well-mixed populations. We consider an evolutionary…
Motivated by the fact that the same social dilemma can be perceived differently by different players, we here study evolutionary multigames in structured populations. While the core game is the weak prisoner's dilemma, a fraction of the…
The application of incentives, such as reward and punishment, is a frequently applied way for promoting cooperation among interacting individuals in structured populations. However, how to properly use the incentives is still a challenging…
Human decision behaviour is quite diverse. In many games humans on average do not achieve maximal payoff and the behaviour of individual players remains inhomogeneous even after playing many rounds. For instance, in repeated prisoner…
We introduce and study an evolutionary complementarity game where in each round a player of population 1 is paired with a member of population 2. The game is symmetric, and each player tries to obtain an advantageous deal, but when one of…
In a society of completely selfish individuals where everybody is only interested in maximizing his own payoff, does any equilibrium exist for the society? John Nash proved more than 50 years ago that an equilibrium always exists such that…
According to the fundamental principle of evolutionary game theory, the more successful strategy in a population should spread. Hence, during a strategy imitation process a player compares its payoff value to the payoff value held by a…
We consider a dynamic collective choice problem where a large number of players are cooperatively choosing between multiple destinations while being influenced by the behavior of the group. For example, in a robotic swarm exploring a new…
Repeated interaction promotes cooperation among rational individuals under the shadow of future, but it is hard to maintain cooperation when a large number of error-prone individuals are involved. One way to construct a cooperative Nash…
When users stand to gain from certain predictions, they are prone to act strategically to obtain favorable predictive outcomes. Whereas most works on strategic classification consider user actions that manifest as feature modifications, we…
Collective foragers, from animals to robotic swarms, must balance exploration and exploitation to locate sparse resources efficiently. While social learning is known to facilitate this balance, how the range of information sharing shapes…
Decision-making individuals are typically either an imitator, who mimics the action of the most successful individual(s), a conformist (or coordinating individual), who chooses an action if enough others have done so, or a nonconformist (or…
The pursuit of highest payoffs in evolutionary social dilemmas is risky and sometimes inferior to conformity. Choosing the most common strategy within the interaction range is safer because it ensures that the payoff of an individual will…
We present Self-Play Preference Optimization (SPO), an algorithm for reinforcement learning from human feedback. Our approach is minimalist in that it does not require training a reward model nor unstable adversarial training and is…