Related papers: MRL order, log-concavity and an application to pea…
Option-critic learning is a general-purpose reinforcement learning (RL) framework that aims to address the issue of long term credit assignment by leveraging temporal abstractions. However, when dealing with extended timescales, discounting…
Many examples of exactly solvable birth and death processes, a typical stationary Markov chain, are presented together with the explicit expressions of the transition probabilities. They are derived by similarity transforming exactly…
Random processes with stationary increments and intrinsic random processes are two concepts commonly used to deal with non-stationary random processes. They are broader classes than stationary random processes and conceptually closely…
A simple model of the new notion of "Markov up" processes is proposed; its positive recurrence and ergodic properties are shown under the appropriate conditions.
Most results regarding Skorokhod embedding problems (SEP) so far rely on the assumption that the corresponding stopped process is uniformly integrable, which is equivalent to the convex ordering condition…
This work is a continuation of [Kalikaeva, MPRF, 23(2):225-240]. The object of study is ``Markov-up processes'' on $\mathbb Z_+$ and the moment of downcrossing a certain barrier. The processes considered in this paper differ from Markov…
Learned representations are a central component in modern ML systems, serving a multitude of downstream tasks. When training such representations, it is often the case that computational and statistical constraints for each downstream task…
We study reinforcement learning by combining recent advances in regularized linear programming formulations with the classical theory of stochastic approximation. Motivated by the challenge of designing algorithms that leverage off-policy…
Symbolic indefinite integration in Computer Algebra Systems such as Maple involves selecting the most effective algorithm from multiple available methods. Not all methods will succeed for a given problem, and when several do, the results,…
Multi-objective reinforcement learning (MORL) is the generalization of standard reinforcement learning (RL) approaches to solve sequential decision making problems that consist of several, possibly conflicting, objectives. Generally, in…
Non-linear Hawkes processes with memory kernels given by the sum of Erlang kernels are considered. It is shown that their stability properties can be studied in terms of an associated class of piecewise deterministic Markov processes,…
We consider a class of semi-Markov processes (SMP) such that the embedded discrete time Markov chain may be non-homogeneous. The corresponding augmented processes are represented as semi-martingales using stochastic integral equation…
We provide a statistical analysis of regularization-based continual learning on a sequence of linear regression tasks, with emphasis on how different regularization terms affect the model performance. We first derive the convergence rate…
We propose trace logic, an instance of many-sorted first-order logic, to automate the partial correctness verification of programs containing loops. Trace logic generalizes semantics of program locations and captures loop semantics by…
We use the abstract method of (local) martingale problems in order to give criteria for convergence of stochastic processes. Extending previous notions, the formulation we use is neither restricted to Markov processes (or semimartingales),…
The proliferation of fast, dense, byte-addressable nonvolatile memory suggests that data might be kept in pointer-rich "in-memory" format across program runs and even process and system crashes. For full generality, such data requires…
We consider mark-recapture-recovery (MRR) data of animals where the model parameters are a function of individual time-varying continuous covariates. For such covariates, the covariate value is unobserved if the corresponding individual is…
We consider Model-Agnostic Meta-Learning (MAML) methods for Reinforcement Learning (RL) problems, where the goal is to find a policy using data from several tasks represented by Markov Decision Processes (MDPs) that can be updated by one…
Recent years have seen a rise in interest in terms of using machine learning, particularly reinforcement learning (RL), for production scheduling problems of varying degrees of complexity. The general approach is to break down the…
We present a new approach to termination analysis of logic programs. The essence of the approach is that we make use of general term-orderings (instead of level mappings), like it is done in transformational approaches to logic program…