Related papers: Exploratory Control with Tsallis Entropy for Laten…
Nonadditive Tsallis $q$-statistics has successfully been applied for a plethora of systems in natural sciences and other branches of knowledge. Nevertheless, its foundations have been severely criticised by some authors based on the…
We propose an open loop methodology based on sample statistics to solve chance constrained stochastic optimal control problems with probabilistic safety guarantees for linear systems where the additive Gaussian noise has unknown mean and…
In this paper, we develop a theoretical framework for nonlinear stochastic optimal control problems with optimal stopping by establishing a density-based deterministic representation of the underlying diffusion. For state-independent…
In actor-critic-based reinforcement learning algorithms such as Twin Delayed Deep Deterministic policy gradient (TD3), insufficient exploration of the spatial space can result in suboptimal policies when controlling 7-DOF robotic arms. To…
The q-Gaussians are discussed from the point of view of variance mixtures of normals and exchangeability. For each q< 3, there is a q-Gaussian distribution that maximizes the Tsallis entropy under suitable constraints. This paper shows that…
We present a variational free-energy formulation for distributionally robust decision-making with ambiguity in the generative model. The formulation, related to a broad range of learning and control frameworks, yields a minimax optimal…
In this paper, we present some geometric properties of the maximum entropy (MaxEnt) Tsallis- distributions under energy constraint. In the case q > 1, these distributions are proved to be marginals of uniform distributions on the sphere; in…
Learning optimal policies in sparse rewards settings is difficult as the learning agent has little to no feedback on the quality of its actions. In these situations, a good strategy is to focus on exploration, hopefully leading to the…
We study the numerical realisation of optimal consensus control laws for agent-based models. For a nonlinear multi-agent system of Cucker-Smale type, consensus control is cast as a dynamic optimisation problem for which we derive…
We present a control model for an octopus tentacle, based on the dynamics of an inextensible string with curvature constraints and curvature controls. We derive the equations of motion together with an appropriate set of boundary…
The non-extensive statistical mechanics has been applied to describe a variety of complex systems with inherent correlations and feedback loops. Here we present a dynamical model based on previously proposed static model exhibiting in the…
We consider a class of reinforcement-learning systems in which the agent follows a behavior policy to explore a discrete state-action space to find an optimal policy while adhering to some restriction on its behavior. Such restriction may…
Quasi-power law ensembles are discussed from the perspective of nonextensive Tsallis distributions characterized by a nonextensive parameter $q$. A number of possible sources of such distributions are presented in more detail. It is further…
This paper studies the continuous-time q-learning (the continuous time counterpart of Q-learing) for Markov switching system under Tsallis entropy regularization. We address the difficulty in traditional RL algorithms where the Tsallis…
The factorization problem of $q$-exponential distribution within nonextensive statistical mechanics is discussed on the basis of Abe's general pseudoadditivity for equilibrium systems. it is argued that the factorization of compound…
In this paper, we investigate an optimal control problem governed by parabolic equations with measure-valued controls over time. We establish the well-posedness of the optimal control problem and derive the first-order optimality condition…
This paper studies an infinite horizon optimal control problem for discrete-time linear system and quadratic criteria, both with random parameters which are independent and identically distributed with respect to time. In this general…
This paper investigates the social optimality of linear quadratic mean field control systems with unmodeled dynamics. The objective of agents is to optimize the social cost, which is the sum of costs of all agents. By variational analysis…
A promising technique for exploration is to maximize the entropy of visited state distribution, i.e., state entropy, by encouraging uniform coverage of visited state space. While it has been effective for an unsupervised setup, it tends to…
Suboptimal methods in optimal control arise due to a limited computational budget, unknown system dynamics, or a short prediction window among other reasons. Although these methods are ubiquitous, their transient performance remains…