Related papers: Positive reinforced generalized time-dependent P\'…
This is the second part of a two-part investigation. We continue the study of a class of balanced urn schemes on balls of two colors (white and black). At each drawing, a sample of size $m\ge 1$ is drawn from the urn and ball addition rules…
This article presents a short and concise description of stochastic approximation algorithms in reinforcement learning of Markov decision processes. The algorithms can also be used as a suboptimal method for partially observed Markov…
Our main result is to prove almost-sure convergence of a stochastic-approximation algorithm defined on the space of measures on a non-compact space. Our motivation is to apply this result to measure-valued P\'olya processes (MVPPs, also…
In the apportionment problem, a fixed number of seats must be distributed among parties in proportion to the number of voters supporting each party. We study a generalization of this setting, in which voters can support multiple parties by…
Competing urns refers to the random experiment where m balls are dropped, randomly and independently, into urns 1,...,n. Formally, we have a random map $\sigma$ from {1,...,m} to {1,...,n} with the $\sigma(i)$'s i.i.d. With $x_j$ the…
We introduce a new model for contagion spread using a network of interacting finite memory two-color P\'{o}lya urns, which we refer to as the finite memory interacting P\'{o}lya contagion network. The urns interact in the sense that the…
We consider a version of the classical Ehrenfest urn model with two urns and two types of balls: regular and heavy. Each ball is selected independently according to a Poisson process having rate $1$ for regular balls and rate…
We present a distributional approach to theoretical analyses of reinforcement learning algorithms for constant step-sizes. We demonstrate its effectiveness by presenting simple and unified proofs of convergence for a variety of…
Consider an urn containing balls labeled with integer values. Define a discrete-time random process by drawing two balls, one at a time and with replacement, and noting the labels. Add a new ball labeled with the sum of the two drawn…
We study first passage statistics of the Polya urn model. In this random process, the urn contains two types of balls. In each step, one ball is drawn randomly from the urn, and subsequently placed back into the urn together with an…
The paper deals with the problem of finding the best alternatives on the basis of pairwise comparisons when these comparisons need not be transitive. In this setting, we study a reinforcement urn model. We prove convergence to the optimal…
A basic experiment in probability theory is drawing without replacement from an urn filled with multiple balls of different colours. Clearly, it is physically impossible to overdraw, that is, to draw more balls from the urn than it…
Reinforcement learning (RL) problems are fundamental in online decision-making and have been instrumental in finding an optimal policy for Markov decision processes (MDPs). Function approximations are usually deployed to handle large or…
Value-function approximation methods that operate in batch mode have foundational importance to reinforcement learning (RL). Finite sample guarantees for these methods often crucially rely on two types of assumptions: (1) mild distribution…
Consider throwing $n$ balls at random into $m$ urns, each ball landing in urn $i$ with probability $p_i$. Let $S$ be the resulting number of singletons, i.e., urns containing just one ball. We give an error bound for the Kolmogorov distance…
We study the phase transition and the critical properties of a nonlinear P\'{o}lya urn, which is a simple binary stochastic process $X(t)\in \{0,1\},t=1,\cdots$ with a feedback mechanism. Let $f$ be a continuous function from the unit…
We study a discrete-time Markov process $X_n\in\mathbb{R}^d$, for which the distribution of the future increments depends only on the relative ranking of its components (descending order by value). We endow the process with a…
The asymptotic behaviour of a generalised P\'olya--Eggenberger urn is well--known to depend on the spectrum of its replacement matrix: If its dominant eigenvalue $r$ is simple and no other eigenvalue is `large' in the sense that its real…
Consider the multicolored urn model where, after every draw, balls of the different colors are added to the urn in a proportion determined by a given stochastic replacement matrix. We consider some special replacement matrices which are not…
Using P\'{o}lya's urn model with negative replacement we introduce a new Bernstein-type operator and we show that the new operator improves upon the known estimates for the classical Bernstein operator. We also provide numerical evidence…