Related papers: Strongly reinforced P\'olya urns with graph-based …
We study the behaviour of a class of edge-reinforced random walks {on $\mathbb{Z}_+$}, with heterogeneous initial weights, where each edge weight can be updated only when the edge is traversed from left to right. We provide a description…
We answer Problem 11.1 of Janson arXiv:1803.04207 on P\'olya urns associated with stable random walk. Our proof use neither martingales nor trees, but an approximation with a differential equation.
We consider a class of reinforcement processes, called WARMs, on tree graphs. These processes involve a parameter $\alpha$ which governs the strength of the reinforcement, and a collection of Poisson processes indexed by the vertices of the…
We study inhomogeneous Bernoulli bond percolation on the graph $G \times \mathbb{Z}$, where $G$ is a connected quasi-transitive graph. The inhomogeneity is introduced through a random region $R$ around the origin axis…
We present the first reinforcement-learning model to self-improve its reward-modulated training implemented through a continuously improving "intuition" neural network. An agent was trained how to play the arcade video game Pong with two…
Reinforcement learning methods have recently been very successful at performing complex sequential tasks like playing Atari games, Go and Poker. These algorithms have outperformed humans in several tasks by learning from scratch, using only…
Reinforcement learning (RL) methods have been shown to be capable of learning intelligent behavior in rich domains. However, this has largely been done in simulated domains without adequate focus on the process of building the simulator. In…
An urn containing specified numbers of balls of distinct ordered colors is considered. A multiple q-Polya urn model is introduced by assuming that the probability of q-drawing a ball of a specific color from the urn varies geometrically,…
Graphs can be used to represent and reason about systems and a variety of metrics have been devised to quantify their global characteristics. However, little is currently known about how to construct a graph or improve an existing one given…
We convert the DeepMind Mathematics Dataset into a reinforcement learning environment by interpreting it as a program synthesis problem. Each action taken in the environment adds an operator or an input into a discrete compute graph. Graphs…
We study strong equilibria in symmetric capacitated cost-sharing games. In these games, a graph with designated source $s$ and sink $t$ is given, and each edge is associated with some cost. Each agent chooses strategically an $s$-$t$ path,…
In repeated-game applications where both the collusive and non-collusive outcomes can be supported as equilibria, researchers must resolve underlying selection questions if theory will be used to understand counterfactual policies. One…
We consider a special case of the generalized P\'{o}lya's urn model introduced by Benaim et al (2013). Given a finite connected graph $G$, place a bin at each vertex. Two bins are called a pair if they share an edge of $G$. At discrete…
Large and diverse datasets have been the cornerstones of many impressive advancements in artificial intelligence. Intelligent creatures, however, learn by interacting with the environment, which changes the input sensory signals and the…
We study a networked system of innovation processes, where each process is modeled as an urn with infinitely many colors-a classical framework for capturing the emergence of novelties. Extending this paradigm, we analyze a model of…
Cumulative advantage (CA) refers to the notion that accumulated resources foster the accumulation of further resources in competitions, a phenomenon that has been empirically observed in various contexts. The oldest and arguably simplest…
Models based on preferential attachment have had much success in reproducing the power law degree distributions which seem ubiquitous in both natural and engineered systems. Here, rather than assuming preferential attachment, we give an…
The P\'olya urn scheme is a discrete-time process concerning the addition and removal of colored balls. There is a known embedding of it in continuous-time, called the P\'olya process. We deal with a generalization of this stochastic model,…
This paper concerns the long term behaviour of a growth model describing a random sequential allocation of particles on a finite cycle graph. The model can be regarded as a reinforced urn model with graph-based interactions. It is motivated…
The problem of reinforcement learning is considered where the environment or the model undergoes a change. An algorithm is proposed that an agent can apply in such a problem to achieve the optimal long-time discounted reward. The algorithm…