Related papers: Learning Expected Reward for Switched Linear Contr…
We propose $\textit{iterative inversion}$ -- an algorithm for learning an inverse function without input-output pairs, but only with samples from the desired output distribution and access to the forward function. The key challenge is a…
We continue our study of the dynamics of mappings with small topological degree on (projective) complex surfaces. Previously, under mild hypotheses, we have constructed an ergodic ``equilibrium'' measure for each such mapping. Here we study…
In many applications, it is often necessary to sample the mean value of certain quantity with respect to a probability measure {\mu} on the level set of a smooth function $\xi: \mathbb{R}^d\rightarrow \mathbb{R}^k$, $1\le k < d$. A…
In this paper, we propose an adaptive event-triggered reinforcement learning control for continuous-time nonlinear systems, subject to bounded uncertainties, characterized by complex interactions. Specifically, the proposed method is…
This paper contains two parts. In the first part, we study the ergodicity of periodic measures of random dynamical systems on a separable Banach space. We obtain that the periodic measure of the continuous time skew-product dynamical system…
We consider reinforcement learning (RL) in episodic Markov decision processes (MDPs) with linear function approximation under drifting environment. Specifically, both the reward and state transition functions can evolve over time but their…
In this article, we study the ergodic risk-sensitive control problem for controlled regime-switching diffusions. Under a blanket stability hypothesis, we solve the associated nonlinear eigenvalue problem for weakly coupled systems and…
We study the problem of system identification for stochastic continuous-time dynamics, based on a single finite-length state trajectory. We present a method for estimating the possibly unstable open-loop matrix by employing properly…
In this paper we consider the problem of learning the optimal policy for uncontrolled restless bandit problems. In an uncontrolled restless bandit problem, there is a finite set of arms, each of which when pulled yields a positive reward.…
We derive consistency and asymptotic normality results for quasi-maximum likelihood methods for drift parameters of ergodic stochastic processes observed in discrete time in an underlying continuous-time setting. The special feature of our…
We consider continuous-time random walk models described by arbitrary sojourn time probability density functions. We find a general expression for the distribution of time-averaged observables for such systems, generalizing some recent…
Via operator theoretic methods, we formalize the concentration phenomenon for a given observable `$r$' of a discrete time Markov chain with `$\mu_{\pi}$' as invariant ergodic measure, possibly having support on an unbounded state space. The…
We present a version of the stochastic maximum principle (SMP) for ergodic control problems. In particular we give necessary (and sufficient) conditions for optimality for controlled dissipative systems in finite dimensions. The strategy we…
We study learning of probability distributions characterized by an unknown symmetry direction. Based on an entropic performance measure and the variational method of statistical mechanics we develop exact upper and lower bounds on the…
In this work, we study non-asymptotic bounds on correlation between two time realizations of stable linear systems with isotropic Gaussian noise. Consequently, via sampling from a sub-trajectory and using \emph{Talagrands'} inequality, we…
We prove existence of (at most denumerable many) absolutely continuous invariant probability measures for random one-dimensional dynamical systems with asymptotic expansion. If the rate of expansion (Lyapunov exponents) is bounded away from…
Learning-based control methods typically assume stationary system dynamics, an assumption often violated in real-world systems due to drift, wear, or changing operating conditions. We study reinforcement learning for control under…
This paper proposes a new adaptation methodology to find the control inputs for a class of nonlinear systems with time-varying bounded uncertainties. The proposed method does not require any prior knowledge of the uncertainties including…
We study multi-objective reinforcement learning with nonlinear preferences over trajectories. That is, we maximize the expected value of a nonlinear function over accumulated rewards (expected scalarized return or ESR) in a multi-objective…
For a large class of transitive non-hyperbolic systems, we construct nonhyperbolic ergodic measures with entropy arbitrarily close to its maximal possible value. The systems we consider are partially hyperbolic with one-dimension central…