Related papers: The discrete renewal theorem with bounded intereve…
In this paper we consider the problem of obtaining sharp bounds for the performance of temporal difference (TD) methods with linear function approximation for policy evaluation in discounted Markov decision processes. We show that a simple…
We prove that if a solution of the discrete time-dependent Schr\"odinger equation with bounded real potential decays fast at two distinct times then the solution is trivial. For the free Shr\"odinger operator and for operators with…
We give a survey of a number of simple applications of renewal theory to problems on random strings and tries: insertion depth, size, insertion mode and imbalance of tries; variations for b-tries and Patricia tries; Khodak and Tunstall…
Temporal-difference (TD) learning is widely regarded as one of the most popular algorithms in reinforcement learning (RL). Despite its widespread use, it has only been recently that researchers have begun to actively study its finite time…
Second order recurrence of a $d$-dimensional diffusion with an additive Wiener process, with switching, and with one recurrent and one transient regime and constant switching intensities is established under suitable conditions. The…
The purpose of this paper is to establish Picard-Lindel\"{o}f theorem for local uniqueness and existence results for first-order systems of nonlinear delay dynamic equations. In the linear case, we extend our results to global existence and…
Let $(\xi_k,\eta_k)_{k\in\mathbb{N}}$ be independent identically distributed random vectors with arbitrarily dependent positive components. We call a (globally) perturbed random walk a random sequence $T:=(T_k)_{k\in\mathbb{N}}$ defined by…
We reformulate and generalize the uniqueness and existence proofs of time-dependent density-functional theory. The central idea is to restate the fundamental one-to-one correspondence between densities and potentials as a global fixed point…
Based on various strategies and a new general doubling operator, we obtain several simple proofs of the celebrated Sharkovsky's cycle coexistence theorem. A simple non-directed graph proof which is especially suitable for a calculus course…
Recently the so-called Prabhakar generalization of the fractional Poisson counting process attracted much interest for his flexibility to adapt real world situations. In this renewal process the waiting times between events are IID…
We analyse quantile temporal-difference learning (QTD), a distributional reinforcement learning algorithm that has proven to be a key component in several successful large-scale applications of reinforcement learning. Despite these…
We extend the theory of d-separation to cases in which data instances are not independent and identically distributed. We show that applying the rules of d-separation directly to the structure of probabilistic models of relational data…
The reconstruction of an unknown function $f$ from its line sums is the aim of discrete tomography. However, two main aspects prevent reconstruction from being an easy task. In general, many solutions are allowed due to the presence of the…
We consider renewal-type processes whose positive inter-renewal times may be dependent, non-identically distributed, and may have mixed distributions. We introduce a generalised intensity measure extending the classical hazard-rate…
In the renewal processes, if the waiting time probability density function is a tempered power-law distribution, then the process displays a transition dynamics; and the transition time depends on the parameter $\lambda$ of the exponential…
We present an algorithm for marginalising changepoints in time-series models that assume a fixed number of unknown changepoints. Our algorithm is differentiable with respect to its inputs, which are the values of latent random variables…
We establish the existence theory of several commonly used finite element (FE) nonlinear fully discrete solutions, and the convergence theory of a linearized iteration. First, it is shown for standard FE, SUPG and edge-averaged method…
Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process. Although TD is one of the most widely used algorithms in reinforcement…
Large deviations for the local time of a process $X_t$ are investigated, where $X_t=x_i$ for $t \in [S_{i-1},S_i[$ and $(x_j)$ are i.i.d.\ random variables on a Polish space, $S_j$ is the $j$-th arrival time of a renewal process depending…
A method for machine learning and serving of discrete field theories in physics is developed. The learning algorithm trains a discrete field theory from a set of observational data on a spacetime lattice, and the serving algorithm uses the…