English
Related papers

Related papers: Positive reinforced generalized time-dependent P\'…

200 papers

Deep Reinforcement Learning has shown significant progress in extracting useful representations from high-dimensional inputs albeit using hand-crafted auxiliary tasks and pseudo rewards. Automatically learning such representations in an…

Machine Learning · Computer Science 2023-06-28 Somjit Nath , Gopeshh Raaj Subbaraj , Khimya Khetarpal , Samira Ebrahimi Kahou

In the classical Polya urn problem, one begins with $d$ bins, each containing one ball. Additional balls arrive one at a time, and the probability that an arriving ball is placed in a given bin is proportional to $m^\gamma$, where $m$ is…

Probability · Mathematics 2014-07-01 Jeremy Chen

In the present paper we show that in P\'{o}lya's urn model, for an arbitrarily fixed initial distribution of the urn, the corresponding random variables satisfy a convex ordering with respect to the replacement parameter. As an application,…

Probability · Mathematics 2022-01-19 Diana Meleşteu , Mihai N. Pascu , Nicolae R. Pascu

We study the problem of Reinforcement Learning (RL) with linear function approximation, i.e. assuming the optimal action-value function is linear in a known $d$-dimensional feature mapping. Unfortunately, however, based on only this…

Machine Learning · Computer Science 2022-11-15 Zeyu Jia , Randy Jia , Dhruv Madeka , Dean P. Foster

A stochastic algorithm is proposed, finding the set of generalized means associated to a probability measure on a compact Riemannian manifold M and a continuous cost function on the product of M by itself. Generalized means include p-means…

Probability · Mathematics 2013-05-28 Marc Arnaudon , Laurent Miclo

Consider a balls-in-bins process in which each new ball goes into a given bin with probability proportional to f(n), where n is the number of balls currently in the bin and f is a fixed positive function. It is known that these so-called…

Probability · Mathematics 2007-07-09 Roberto Imbuzeiro Oliveira

We introduce and study a discrete multi-period extension of the classical knapsack problem, dubbed generalized incremental knapsack. In this setting, we are given a set of $n$ items, each associated with a non-negative weight, and $T$ time…

Data Structures and Algorithms · Computer Science 2020-09-16 Yuri Faenza , Danny Segev , Lingyi Zhang

In a recent article a generalization of the binomial distribution associated with a sequence of positive numbers was examined. The analysis of the nonnegativeness of the formal expressions was a key-point to allow to give them a statistical…

Mathematical Physics · Physics 2015-06-04 H. Bergeron , E. M. F. Curado , J. P. Gazeau , Ligia M. C. S. Rodrigues

Normalised generalised gamma processes are random probability measures that induce nonparametric prior distributions widely used in Bayesian statistics, particularly for mixture modelling. We construct a class of dependent normalised…

Probability · Mathematics 2016-11-07 Matteo Ruggiero , Matteo Sordello

We consider systems of stochastic fixed-point equations that arise in the asymptotic analysis of random recursive structures and algorithms such as Quicksort, generalized P\'olya urn processes and path lengths of random recursive trees and…

Probability · Mathematics 2018-03-08 Kevin Leckey

We consider Reinforced Random Walks where transition probabilities are a function of the proportion of times the walk has traversed an edge. We give conditions for recurrence or transience. A phase transition is observed, similar to…

Probability · Mathematics 2009-07-15 Olivier Raimond , Bruno Schapira

The Generalized P\'{o}lya Urn (GPU) is a popular urn model which is widely used in many disciplines. In particular, it is extensively used in treatment allocation schemes in clinical trials. In this paper, we propose a sequential…

Probability · Mathematics 2007-05-23 Li-X. Zhang , Feifang Hu , Siu Hung Cheung

We define the notions of disjoint unions and products for generalised P\'olya urns, proving that this turns the set of isomorphism classes of urns into a commutative semiring. The set of square matrices up to similarity by a permutation…

Probability · Mathematics 2021-11-16 Fabian Burghart

This paper presents the first sufficient conditions that guarantee the stability and almost sure convergence of multi-timescale stochastic approximation (SA) iterates. It extends the existing results on one-timescale and two-timescale SA…

Systems and Control · Electrical Eng. & Systems 2025-10-16 Rohan Deb , Swetha Ganesh , Shalabh Bhatnagar

Consider an urn model whose replacement matrix is triangular, has all entries nonnegative and the row sums are all equal to one. We obtain the strong laws for the counts of balls corresponding to each color. The scalings for these laws…

Probability · Mathematics 2010-09-27 Arup Bose , Amites Dasgupta , Krishanu Maulik

For the renormalised sums of the random $\pm 1$-colouring of the connected components of $\mathbb Z$ generated by the coalescing renewal processes in the "power law P\'olya's urn" of Hammond and Sheffield we prove functional convergence…

Probability · Mathematics 2023-01-18 Jan Lukas Igelbrink , Anton Wakolbinger

The random vector of frequencies in a generalized urn model is viewed as conditionally independent random variables, given their sum. Such a representation is exploited to derive Edgeworth expansions for a sum of functions of such…

Probability · Mathematics 2014-01-20 Sh. M. Mirakhmedov , S. Rao Jammalamadaka , Ibrahim B. Mohamed

Consider a probability measure supported by a regular geodesic ball in a manifold. For any p larger than or equal to 1 we define a stochastic algorithm which converges almost surely to the p-mean of the measure. Assuming furthermore that…

Probability · Mathematics 2011-06-28 Marc Arnaudon , Clément Dombry , Anthony Phan , Le Yang

Reinforcement learning has been applied to many interesting problems such as the famous TD-gammon and the inverted helicopter flight. However, little effort has been put into developing methods to learn policies for complex persistent tasks…

Artificial Intelligence · Computer Science 2016-06-22 Xiao Li , Calin Belta

We consider a preferential attachment random graph with self-reinforcement. Each time a new vertex comes in, it attaches itself to an old vertex with a probability that is proportional to the sum of the degrees of that old vertex at all…

Probability · Mathematics 2025-07-29 Yogesh Dahiya , Frank den Hollander