Related papers: Renewal processes with costs and rewards
It is more important to estimate the rate of convergence to a stationary distribution rather than only to prove the existence one in many applied problems of reliability and queuing theory. This can be done via standard methods, but only…
In repeated interactions between individuals, we do not expect that exactly the same situation will occur from one time to another. Contrary to what is common in models of repeated games in the literature, most real situations may differ a…
The choice titration procedure presents a subject with a repeated choice between a standard option that always provides the same reward and an adjusting option for which the reward schedule is adjusted based on the subjects previous…
Model-checking resource logics with production and consumption of resources is a computationally hard and often undecidable problem. We introduce a simple and realistic assumption that there is at least one diminishing resource, that is, a…
We propose analyzing conditional reasoning by appeal to a notion of intervention on a simulation program, formalizing and subsuming a number of approaches to conditional thinking in the recent AI literature. Our main results include a…
In this note we explain two transitions known for moment generating functions of local times by means of properties of the renewal measure of a related renewal equation. The arguments simplify and strengthen results on the asymptotic…
While reinforcement learning (RL) demonstrated remarkable success in enhancing the reasoning capabilities of language models, the training dynamics of RL in LLMs remain unclear. In this work, we provide an explanation of the RL training…
This paper investigates the so-called reward-balancing methods, a novel class of algorithms for solving discounted-return reinforcement learning (RL) problems. These methods consist of iteratively adjusting the reward function to transform…
Let $\{q_n\}_{n=0}^\infty\subset [0,1]$ satisfy $q_0=0$, $\sum_{n=0}^\infty q_n=1$, and $\gcd\{n\geq 1\mid q_n\neq 0\}=1$. We consider the following process: Let $x$ be a real number. We first set $x=0$. Then $x$ is increased by $i$ with…
The paper deals with renewal theory for a class of extremal Markov sequences connected with the Kendall convolution. We consider here some particular cases of the Wold processes associated with generalized convolutions. We prove an analogue…
In this short note, we give two proofs of the infinitude of primes via valuation theory and give a new proof of the divergence of the sum of prime reciprocals by Roth's theorem and Euler-Legendre's theorem for arithmetic progressions.
Sophisticated machine models are increasingly used for high-stakes decisions in everyday life. There is an urgent need to develop effective explanation techniques for such automated decisions. Rule-Based Explanations have been proposed for…
Inverse reinforcement learning (IRL) denotes a powerful family of algorithms for recovering a reward function justifying the behavior demonstrated by an expert agent. A well-known limitation of IRL is the ambiguity in the choice of the…
We consider a renewal process which models a cumulative shock model that fails when the accumulation of shocks up-crosses a certain threshold. The ratio limit properties of the probabilities of non-failure after n cumulative shocks are…
Combining a standard proof search method, such as resolution or tableaux, and rewriting is a powerful way to cut off search space in automated theorem proving, but proving the completeness of such combined methods may be challenging. It may…
We extend the super-replication theorems of [27] in a dynamic setting, both in the num\'eraire-based as well as in the num\'eraire-free setting. For this purpose, we generalize the notion of admissible strategies. In particular, we obtain a…
This paper considers an opportunistic scheduling problem over a renewal system. A controller observes a random event at the beginning of each renewal frame and then chooses an action in response to the event, which affects the duration of…
If the universe follows a specific design, then a central question is which cost function is optimized by the observed forces. This is the problem of inverse optimal control, or inverse reinforcement learning, in which a reward function is…
We study the set of (stationary) feasible payoffs of overlapping generation repeated games that can be achieved by action sequences in which every generation of players plays the same sequence of action profiles. First, we completely…
Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that the reasoning chains trained through RLVR reliably represent…