Related papers: Exploratory Control with Tsallis Entropy for Laten…
We give a new proof of the theorems on the maximum entropy principle in Tsallis statistics. That is, we show that the $q$-canonical distribution attains the maximum value of the Tsallis entropy, subject to the constraint on the…
Exploration is essential for reinforcement learning (RL). To face the challenges of exploration, we consider a reward-free RL framework that completely separates exploration from exploitation and brings new challenges for exploration…
We establish a connection between stochastic optimal control and generative models based on stochastic differential equations (SDEs), such as recently developed diffusion probabilistic models. In particular, we derive a…
Reinforcement Learning has drawn huge interest as a tool for solving optimal control problems. Solving a given problem (task or environment) involves converging towards an optimal policy. However, there might exist multiple optimal policies…
This paper considers the problem of designing time-dependent, real-time control policies for controllable nonlinear diffusion processes, with the goal of obtaining maximally-informative observations about parameters of interest. More…
In this work, we study the control constrained distributed optimal control of a stationary doubly diffusive flow model. For the control problem, we use a well-posedness analysis based on minimal assumptions on data and domain. We show the…
The focus of this paper is directed towards optimal control of multi-agent systems consisting of one leader and a number of followers in the presence of noise. The dynamics of every agent is assumed to be linear, and the performance index…
Efficient exploration remains a central challenge in reinforcement learning, serving as a useful pretraining objective for data collection, particularly when an external reward function is unavailable. A principled formulation of the…
A finite horizon linear quadratic(LQ) optimal control problem is studied for a class of discrete-time linear fractional systems (LFSs) affected by multiplicative, independent random perturbations. Based on the dynamic programming technique,…
Financial markets are highly non-linear and non-equilibrium systems. Earlier works have suggested that the behavior of market returns can be well described within the framework of non-extensive Tsallis statistics or superstatistics. For…
We consider the problem of stochastic optimal control, where the state-feedback control policies take the form of a probability distribution and where a penalty on the entropy is added. By viewing the cost function as a Kullback- Leibler…
In this paper, we present a new class of Markov decision processes (MDPs), called Tsallis MDPs, with Tsallis entropy maximization, which generalizes existing maximum entropy reinforcement learning (RL). A Tsallis MDP provides a unified…
This paper investigates the optimal control problem for a class of parabolic equations where the diffusion coefficient is influenced by a control function acting nonlocally. Specifically, we consider the optimization of a cost functional…
In this paper, we propose a novel maximum causal Tsallis entropy (MCTE) framework for imitation learning which can efficiently learn a sparse multi-modal policy distribution from demonstrations. We provide the full mathematical analysis of…
The optimization problems defining meta-stable or stationary equilibrium are explored. The Gibbs scheme is modified aiming to describe the statistical properties of a class of non-equilibrium and metastable states. The system is assumed to…
Learning-based techniques are increasingly effective at controlling complex systems using data-driven models. However, most work done so far has focused on learning individual tasks or control laws. Hence, it is still a largely unaddressed…
Gibbs-Boltzmann entropy leads to systems that have a strong dependence on initial conditions. In reality, most materials behave quite independently of initial conditions. Nonextensive entropy or Tsallis entropy leads to nonextensive…
Time distributed optimization is an implementation strategy that can significantly reduce the computational burden of model predictive control by exploiting its robustness to incomplete optimization. When using this strategy, optimization…
Agent behavior is arguably the greatest source of uncertainty in trajectory planning for autonomous vehicles. This problem has motivated significant amounts of work in the behavior prediction community on learning rich distributions of the…
This paper studies social optimal control of mean field LQG (linear-quadratic-Gaussian) models with uncertainty. Specially, the uncertainty is represented by a uncertain drift which is common for all agents. A robust optimization approach…