Related papers: Tsallis Entropy Regularization for Linearly Solvab…
Despite the many recent advances in reinforcement learning (RL), the question of learning policies that robustly satisfy state constraints under unknown disturbances remains open. In this paper, we offer a new perspective on achieving…
We revisit the cut-off prescriptions which are needed in order to specify completely the form of Tsallis' maximum entropy distributions. For values of the Tsallis entropic parameter $q>1$ we advance an alternative cut-off prescription and…
This paper investigates applicability of thermodynamic concepts and principles to competitive systems. We show that Tsallis entropies are suitable for characterisation of systems with transitive competition when mutations deviate from Gibbs…
Tsallis and R\'{e}nyi entropy measures are two possible different generalizations of the Boltzmann-Gibbs entropy (or Shannon's information) but are not generalizations of each others. It is however the Sharma-Mittal measure, which was…
Entropy regularization is an important idea in reinforcement learning, with great success in recent algorithms like Soft Q Network (SQN) and Soft Actor-Critic (SAC1). In this work, we extend this idea into the on-policy realm. We propose…
The entropy of a closure operator has been recently proposed for the study of network coding and secret sharing. In this paper, we study closure operators in relation to their entropy. We first introduce four different kinds of rank…
We present a technique for entropy optimization to calculate a distribution from its moments. The technique is based upon maximizing a discretized form of the Shannon entropy functional by mapping the problem onto a dual space where an…
Entropy is a key measure in studies related to information theory and its many applications. Campbell of the first time recognized that exponential of Shannons entropy is just the size of the sample space when the distribution is uniform.…
The exact solution of a particular form of the stationary state generalized Fokker-Planck equations, which is given under certain conditions by the classical Tsallis distribution, is compared with the solution of the MAXENT equations…
State entropy regularization has empirically shown better exploration and sample complexity in reinforcement learning (RL). However, its theoretical guarantees have not been studied. In this paper, we show that state entropy regularization…
Gauss' law of error is generalized in Tsallis statistics such as multifractal systems, in which Tsallis entropy plays an essential role instead of Shannon entropy. For the generalization, we apply the new multiplication operation determined…
It is possible to derive the maximum entropy principle from thermodynamic stability requirements. Using as a starting point the equilibrium probability distribution, currently used in non-extensive thermostatistics, it turns out that the…
The generalisation and robustness properties of policies learnt through Maximum-Entropy Reinforcement Learning are investigated on chaotic dynamical systems with Gaussian noise on the observable. First, the robustness under noise…
We study the nonextensive thermodynamics for open systems. On the basis of the maximum entropy principle, the dual power-law q-distribution functions are re-deduced by using the dual particle number definitions and assuming that the…
This work uses the entropy-regularised relaxed stochastic control perspective as a principled framework for designing reinforcement learning (RL) algorithms. Herein agent interacts with the environment by generating noisy controls…
Stochastic and soft optimal policies resulting from entropy-regularized Markov decision processes (ER-MDP) are desirable for exploration and imitation learning applications. Motivated by the fact that such policies are sensitive with…
Entropy regularization is an efficient technique for encouraging exploration and preventing a premature convergence of (vanilla) policy gradient methods in reinforcement learning (RL). However, the theoretical understanding of…
The paper extends the analysis of the entropies of the Poisson distribution with parameter $\lambda$. It demonstrates that the Tsallis and Sharma-Mittal entropies exhibit monotonic behavior with respect to $\lambda$, whereas two generalized…
In this work, we present a novel characterization of approximate Nash equilibria in a class of convex games over the simplex. To achieve this, we regularize the utility functions using the Shannon entropy term, connect the solutions to the…
Reinforcement learning (RL) has become a key approach for enhancing reasoning in large language models (LLMs), yet scalable training is often hindered by the rapid collapse of policy entropy, which leads to premature convergence and…