Related papers: Continuous-time q-Learning for Jump-Diffusion Mode…
In recent years there has been a collective research effort to find new formulations of reinforcement learning that are simultaneously more efficient and more amenable to analysis. This paper concerns one approach that builds on the linear…
Quasi-power law ensembles are discussed from the perspective of nonextensive Tsallis distributions characterized by a nonextensive parameter $q$. A number of possible sources of such distributions are presented in more detail. It is further…
The optimal execution problem has always been a continuously focused research issue, and many reinforcement learning (RL) algorithms have been studied. In this article, we consider the execution problem of targeting the volume weighted…
Diffusion models excel at capturing complex data distributions, such as those of natural images and proteins. While diffusion models are trained to represent the distribution in the training dataset, we often are more concerned with other…
Applying Q-learning to high-dimensional or continuous action spaces can be difficult due to the required maximization over the set of possible actions. Motivated by techniques from amortized inference, we replace the expensive maximization…
This paper introduces and analyzes an improved Q-learning algorithm for discrete-time linear time-invariant systems. The proposed method does not require any knowledge of the system dynamics, and it enjoys significant efficiency advantages…
In the present work, we have found that the phenomenological Tsallis distribution (which nowadays is largely used to describe the transverse momentum distributions of hadrons measured in $pp$ collisions at high energies) is consistent with…
The aim of this paper is to investigate the q -> 1/q duality in an information-entropy theory of all q-generalized entropy functionals (Tsallis, Renyi and Sharma-Mittal measures) in the light of a representation based on generalized…
We provide an update of the overview of imprints of Tsallis nonextensive statistics seen in a multiparticle production processes. They reveal an ubiquitous presence of power law distributions of different variables characterized by the…
Recent developments in Reinforcement learning have significantly enhanced sequential decision-making in uncertain environments. Despite their strong performance guarantees, most existing work has focused primarily on improving the…
Maximum entropy models provide the least constrained probability distributions that reproduce statistical properties of experimental datasets. In this work we characterize the learning dynamics that maximizes the log-likelihood in the case…
We explore reinforcement learning methods for finding the optimal policy in the linear quadratic regulator (LQR) problem. In particular, we consider the convergence of policy gradient methods in the setting of known and unknown parameters.…
We show that within classical statistical mechanics it is possible to naturally derive power law distributions which are of Tsallis type. The only assumption is that microcanonical distributions have to be separable from of the total system…
In the past few years, off-policy reinforcement learning methods have shown promising results in their application for robot control. Deep Q-learning, however, still suffers from poor data-efficiency and is susceptible to stochasticity in…
In classical Q-learning, the objective is to maximize the sum of discounted rewards through iteratively using the Bellman equation as an update, in an attempt to estimate the action value function of the optimal policy. Conventionally, the…
A well-known result across information theory, machine learning, and statistical physics shows that the maximum entropy distribution under a mean constraint has an exponential form called the Gibbs-Boltzmann distribution. This is used for…
We present Q-chunking, a simple yet effective recipe for improving reinforcement learning (RL) algorithms for long-horizon, sparse-reward tasks. Our recipe is designed for the offline-to-online RL setting, where the goal is to leverage an…
We discuss the idea that the Tsallis-type (q-additive) entropic chain rule allows for a wider class of entropic functionals than previously thought. In particular, we point out that the ensuing entropy solutions (e.g., Tsallis entropy) can…
The connection between Tsallis entropy for a multifractal distribution and Jackson's $q$-derivative is established. Based on this derivation and definition of a homogeneous function, a $q$-analogue of Shannon's entropy is discussed.…
Continuous-time reinforcement learning offers an appealing formalism for describing control problems in which the passage of time is not naturally divided into discrete increments. Here we consider the problem of predicting the distribution…