Related papers: Renewal processes with costs and rewards
This paper presents a new proof of the renewal theorem by bijecting a general point process to a deterministic one (where the time between events is always fixed). It also provides insight into the workings of the renewal theorem.
In queueing theory, Lorden's inequality can be used for bounds estimation of the moments of backward and forward renewal times. Two random variables called backwards renewal time and forward renewal time for this process are defined.…
Martingales constitute a basic tool in stochastic analysis; this paper considers their application to counting processes. We use this tool to revisit a renewal theorem and its extensions for various counting processes. We first consider a…
We study a renewal problem within a periodic environment, departing from the classical renewal theory by relaxing the assumption of independent and identically distributed inter-arrival times. Instead, the conditional distribution of the…
Exploration in reinforcement learning (RL) remains an open challenge. RL algorithms rely on observing rewards to train the agent, and if informative rewards are sparse the agent learns slowly or may not learn at all. To improve exploration…
Process Reward Models (PRMs) have emerged as a promising approach to enhance the reasoning capabilities of large language models (LLMs) by guiding their step-by-step reasoning toward a final answer. However, existing PRMs either treat each…
A uniform key renewal theorem is deduced from the uniform Blackwell's renewal theorem. A uniform LDP (large deviations principle) for renewal-reward processes is obtained, and MDP (moderate deviations principle) is deduced under conditions…
In this paper we introduce and study renewal-reward processes in random environments where each renewal involves a reward taking values in a Banach space. We derive quenched large deviation principles and identify the associated rate…
This paper considers a particular renewal-reward process with multivariate discounted rewards (inputs) where the arrival epochs are adjusted by adding some random delays. Then this accumulated reward can be regarded as multivariate…
The ability to estimate the rate of convergence for the distributions of regenerative processes is in great demand. These processes are often encountered in queuing theory and in related problems. In some papers on regenerative processes,…
In this paper we investigate the normal and the large fluctuations of additive functionals associated with a stochastic process under a general non-Poissonian resetting mechanism. Cumulative functionals of regenerative processes are very…
Reward models have been increasingly critical for improving the reasoning capability of LLMs. Existing research has shown that a well-trained reward model can substantially improve model performances at inference time via search. However,…
The reward hypothesis posits that, "all of what we mean by goals and purposes can be well thought of as maximization of the expected value of the cumulative sum of a received scalar signal (reward)." We aim to fully settle this hypothesis.…
The (extended) AGM postulates for belief revision seem to deal with the revision of a given theory K by an arbitrary formula, but not to constrain the revisions of two different theories by the same formula. A new postulate is proposed and…
We evaluate the average waiting time between observing the price of financial markets and the next price change, especially in an on-line foreign exchange trading service for individual customers via the internet. Basic technical idea of…
In this paper, we build a new test of rational expectations based on the marginal distributions of realizations and subjective beliefs. This test is widely applicable, including in the common situation where realizations and beliefs are…
If the step distribution in a renewal process has finite mean and regularly varying tail with index -{\alpha}, 1<{\alpha}<2, the first two terms in the asymptotic expansion of the renewal function have been known for many years. Here we…
We give a survey of a number of simple applications of renewal theory to problems on random strings and tries: insertion depth, size, insertion mode and imbalance of tries; variations for b-tries and Patricia tries; Khodak and Tunstall…
Renewal processes are broadly used to model stochastic behavior consisting of isolated events separated by periods of quiescence, whose durations are specified by a given probability law. Here, we identify the minimal sufficient statistic…
Markov decision processes are useful models of concurrency optimisation problems, but are often intractable for exhaustive verification methods. Recent work has introduced lightweight approximative techniques that sample directly from…