Related papers: MRL order, log-concavity and an application to pea…
The analytic form of a new class of factorized Runge-Kutta-Chebyshev (FRKC) stability polynomials of arbitrary order $N$ is presented. Roots of FRKC stability polynomials of degree $L=MN$ are used to construct explicit schemes comprising…
Self-paced reinforcement learning (RL) aims to improve the data efficiency of learning by automatically creating sequences, namely curricula, of probability distributions over contexts. However, existing techniques for self-paced RL fail in…
We consider a general honest homogeneous continuous-time Markov process with restarts. The process is forced to restart from a given distribution at time moments generated by an independent Poisson process. The motivation to study such…
Entropy regularized Markov decision processes have been widely used in reinforcement learning. This paper is concerned with the primal-dual formulation of the entropy regularized problems. Standard first-order methods suffer from slow…
Stochastic and soft optimal policies resulting from entropy-regularized Markov decision processes (ER-MDP) are desirable for exploration and imitation learning applications. Motivated by the fact that such policies are sensitive with…
Mean-field models approximate large stochastic systems by simpler differential equations that are supposed to approximate the mean of the larger system. It is generally assumed that as the stochastic systems get larger (i.e., more people or…
Multi-objective reinforcement learning (MORL) approaches have emerged to tackle many real-world problems with multiple conflicting objectives by maximizing a joint objective function weighted by a preference vector. These approaches find…
Under a fourth order moment condition on the branching and a second order moment condition on the immigration mechanisms, we show that an appropriately scaled projection of a supercritical and irreducible continuous state and continuous…
Higher-order recursion schemes are a higher-order analogue of Boolean Programs; they form a natural class of abstractions for functional programs. We present a new, efficient algorithm for checking CTL properties of the trees generated by…
We study a generalization of the recently introduced order-preserving pattern matching, where instead of looking for an exact copy of the pattern, we only require that the relative order between the elements is the same. In our variant, we…
Monotone L\'evy processes with additive increments are defined and studied. It is shown that these processes have a natural Markov structure and their Markov transition semigroups are characterized using the monotone L\'evy-Khintchine…
Due to recent breakthroughs, reinforcement learning (RL) has demonstrated impressive performance in challenging sequential decision-making problems. However, an open question is how to make RL cope with partial observability which is…
In simulation-based inferences for partially observed Markov process models (POMP), the by-product of the Monte Carlo filtering is an approximation of the log likelihood function. Recently, iterated filtering [14, 13] has originally been…
We determine sufficient conditions under which certain recursively defined functions are well defined for all real inputs. Given a function $f:\mathbb R\to\mathbb R$, call a decreasing sequence $x_1>x_2>x_3>\cdots$ "$f$-bad" if…
The multi-scale entanglement renormalisation ansatz (MERA) is argued to provide a natural description for topological states of matter. The case of Kitaev's toric code is analyzed in detail and shown to possess a remarkably simple MERA…
Large language models (LLMs) have shown promise in performing complex multi-step reasoning, yet they continue to struggle with mathematical reasoning, often making systematic errors. A promising solution is reinforcement learning (RL)…
We survey some of the mechanisms used to prove that naturally defined sequences in combinatorics are log-concave. Among these mechanisms are Alexandrov's inequality for mixed discriminants, the Alexandrov Fenchel inequality for mixed…
Markov decision processes are typically used for sequential decision making under uncertainty. For many aspects however, ranging from constrained or safe specifications to various kinds of temporal (non-Markovian) dependencies in task and…
We present a new approach to termination analysis of logic programs. The essence of the approach is that we make use of general orderings (instead of level mappings), like it is done in transformational approaches to logic program…
We show that any stochastically monotone Feller semigroup on R can be extended by a consistent family of order-preserving Feller semigroups on the successive powers of R. We exhibit a specific such family, which is uniquely characterized by…