Related papers: On Occupation Time for On-Off Processes with Multi…
This paper develops a decision algorithm for weak bisimulation on Markov Automata (MA). For that purpose, different notions of vanishing state (a concept known from the area of Generalised Stochastic Petri Nets) are defined. Vanishing…
Off-policy learning is a framework for evaluating and optimizing policies without deploying them, from data collected by another policy. Real-world environments are typically non-stationary and the offline learned policies should adapt to…
For a network of discrete states with a periodically driven Markovian dynamics, we develop an inference scheme for an external observer who has access to some transitions. Based on waiting-time distributions between these transitions, the…
Tied-down renewal processes are generalisations of the Brownian bridge, where an event (or a zero crossing) occurs both at the origin of time and at the final observation time $t$. We give an analytical derivation of the two-time…
We study offline multi-agent reinforcement learning (RL) in Markov games, where the goal is to learn an approximate equilibrium -- such as Nash equilibrium and (Coarse) Correlated Equilibrium -- from an offline dataset pre-collected from…
Occupied diffusions offer a Markovian framework for path-dependent dynamics by lifting the state space with a flow of occupation measures. Because this additional feature is infinite-dimensional, the simulation of these processes remains…
This paper deals with control of partially observable discrete-time stochastic systems. It introduces and studies Markov Decision Processes with Incomplete Information and with semi-uniform Feller transition probabilities. The important…
The NP-hard problem of task scheduling with communication delays (P|prec,c_{ij}|C_{\mathrm{max}}) is often tackled using approximate methods, but guarantees on the quality of these heuristic solutions are hard to come by. Optimal schedules…
In this paper, we study the problem of continuous-time state observation over lossy communication networks. We consider the situation in which the samplers for measuring the output of the plant are spatially distributed and their…
We study the existence of densities for distributions of piecewise deterministic Markov processes. We also obtain relationships between invariant densities of the continuous time process and that of the process observed at jump times. In…
Long-time limit of one-dimensional L\'{e}vy processes weighted and normalized with respect to the exponential functional of two-point local times are studied. The limit processes may vary according to the choice of random clocks.
We investigate under which conditions a single simulation of joint default times at a final time horizon can be decomposed into a set of simulations of joint defaults on subsequent adjacent sub-periods leading to that final horizon. Besides…
Optimization of decision problems in stochastic environments is usually concerned with maximizing the probability of achieving the goal and minimizing the expected episode length. For interacting agents in time-critical applications,…
We introduce a novel framework for analyzing reinforcement learning (RL) in continuous state-action spaces, and use it to prove fast rates of convergence in both off-line and on-line settings. Our analysis highlights two key stability…
We consider a time-slotted job-assignment system with a central server, N users and a machine which changes its state according to a Markov chain (hence called a Markov machine). The users submit their jobs to the central server according…
This paper presents a new condition for the existence of optimal stationary policies in average-cost continuous-time Markov decision processes with unbounded cost and transition rates, arising from controlled queueing systems. This…
We find the moment generating function (mgf) of the nonequilibrium work for open systems undergoing a thermal process, ie, when the stochastic dynamics maps thermal states into time dependent thermal states. The mgf is given in terms of a…
We consider a time-slotted communication system with a machine, a cloud server, and a sampler. Job requests from the users are queued on the server to be completed by the machine. The machine has two states, namely, a busy state and a free…
Levy walk is a fundamental model with applications ranging from quantum physics to paths of animal foraging. Taking animal foraging as an example, a natural idea that comes to one's mind is to introduce the multiple internal states for…
The theory of ``Markov-up'' processes is being developed. This is a new class of stochastic processes with ``partial'' markovian features; it could also be called ``one-sided Markov''. Such a behavior may be found in the real world and in…