Related papers: Exploratory Control with Tsallis Entropy for Laten…
Inference-time scaling (ITS) in latent reasoning models typically relies on heuristic perturbations, such as dropout or fixed Gaussian noise, to generate diverse candidate trajectories. However, we show that stronger perturbations do not…
The problem of pure exploration in Markov decision processes has been cast as maximizing the entropy over the state distribution induced by the agent's policy, an objective that has been extensively studied. However, little attention has…
This paper explores a class of fully coupled nonlinear forward-backward stochastic difference equations (FBS$\Delta$Es). Building on insights from linear quadratic optimal control problems, we introduce a more relaxed framework of…
In this paper, we provide a generalized framework for Variational Inference-Stochastic Optimal Control by using thenon-extensive Tsallis divergence. By incorporating the deformed exponential function into the optimality likelihood function,…
The behavior of stock market returns over a period of 1-60 days has been investigated for S&P 500 and Nasdaq within the framework of nonextensive Tsallis statistics. Even for such long terms, the distributions of the returns are…
A policy in deep reinforcement learning (RL), either deterministic or stochastic, is commonly parameterized as a Gaussian distribution alone, limiting the learned behavior to be unimodal. However, the nature of many practical…
We propose an effective exponential model of delay discounting considering fluctuation in impulsivity. This model is seen to be dual to the two-parameter Tsallis model of delay discounting proposed by Takahashi in 2007. We demonstrate that…
We present a computational framework for synthesis of distributed control strategies for a heterogeneous team of robots in a partially observable environment. The goal is to cooperatively satisfy specifications given as Truncated Linear…
In this paper, we study the maximum principle for stochastic optimal control problems of forward-backward stochastic difference systems (FBS{\Delta}Ss). Two types of FBS{\Delta}Ss are investigated. The first one is described by a partially…
In this paper, the authors study the distributed optimal control of a system of three evolutionary equations involving fractional powers of three selfadjoint, monotone, unbounded linear operators having compact resolvents. The system is a…
Animals have a developed ability to explore that aids them in important tasks such as locating food, exploring for shelter, and finding misplaced items. These exploration skills necessarily track where they have been so that they can plan…
The q-exponential distributions, which are generalizations of the Zipf-Mandelbrot power-law distribution, are frequently encountered in complex systems at their stationary states. From the viewpoint of the principle of maximum entropy, they…
This paper investigates an optimal consumption-investment problem featuring recursive utility via Tsallis relative entropy. We establish a fundamental connection between this optimization problem and a quadratic backward stochastic…
We consider a one dimensional elliptic distributed optimal control problem with pointwise constraints on the derivative of the state. By exploiting the variational inequality satisfied by the derivative of the optimal state, we obtain…
A new method is proposed for analyzing complexity and studying the information in random geometric networks using Tsallis entropy tool. Tsallis entropy of the ensemble of random geometric networks is calculated based on the components of…
We study high-dimensional stochastic optimal control problems in which many agents cooperate to minimize a convex cost functional. We consider both the full-information problem, in which each agent observes the states of all other agents,…
We developed a strategic of optimal portfolio based on information theory and Tsallis statistics. The growth rate of a stock market is defined by using $q$-deformed functions and we find that the wealth after n days with the optimal…
In this paper, we study the optimal dividend problem under the continuous time diffusion model with the bounded dividend rate from the Reinforcement Learning (RL) perspective. Unlike the standard literature, our main focus will be on…
We study the problem of exploration in Reinforcement Learning and present a novel model-free solution. We adopt an information-theoretical viewpoint and start from the instance-specific lower bound of the number of samples that have to be…
In this work, we present numerical analysis for a distributed optimal control problem, with box constraint on the control, governed by a subdiffusion equation which involves a fractional derivative of order $\alpha\in(0,1)$ in time. The…