Related papers: Maximum Relative Divergence Principle for Grading …
We consider the minimum error entropy (MEE) criterion and an empirical risk minimization learning algorithm in a regression setting. A learning theory approach is presented for this MEE algorithm and explicit error bounds are provided in…
In ordinary statistical mechanics the Boltzmann-Shannon entropy is related to the Maxwell-Bolzmann distribution $p_i$ by means of a twofold link. The first link is differential and is offered by the Jaynes Maximum Entropy Principle. The…
This article provides a completion to theories of information based on entropy, resolving a longstanding question in its axiomatization as proposed by Shannon and pursued by Jaynes. We show that Shannon's entropy function has a…
Regularization by the Shannon entropy enables us to efficiently and approximately solve optimal transport problems on a finite set. This paper is concerned with regularized optimal transport problems via Bregman divergence. We introduce the…
Entropy notions for $\varepsilon$-incremental practical stability and incremental stability of deterministic nonlinear systems under disturbances are introduced. The entropy notions are constructed via a set of points in state space which…
We describe an approach to improving model fitting and model generalization that considers the entropy of distributions of modelling residuals. We use simple simulations to demonstrate the observational signatures of overfitting on ordered…
Neural networks have dramatically increased our capacity to learn from large, high-dimensional datasets across innumerable disciplines. However, their decisions are not easily interpretable, their computational costs are high, and building…
There are three ways to conceptualize entropy: entropy as an extensive thermodynamic quantity of physical systems (Clausius, Boltzmann, Gibbs), entropy as a measure for information production of ergodic sources (Shannon), and entropy as a…
It is supposed that the exponential multiplier in the method of the non-equilibrium statistical operator (Zubarev`s approach) can be considered as a distribution density of the past lifetime of the system, and can be replaced by an…
This thesis develops a new divergence that generalizes relative entropy and can be used to compare probability measures without a requirement of absolute continuity. We establish properties of the divergence, and in particular derive and…
Product probability property, known in the literature as statistical independence, is examined first. Then generalized entropies are introduced, all of which give generalizations to Shannon entropy. It is shown that the nature of the…
The classical Density Functional Theory (DFT) is introduced as an application of entropic inference for inhomogeneous fluids at thermal equilibrium. It is shown that entropic inference reproduces the variational principle of DFT when…
We derive an upper bound for the mean of the supremum of the empirical process indexed by a class of functions that are known to have variance bounded by a small constant $\delta$. The bound is expressed in the uniform entropy integral of…
Standard Markov decision process (MDP) and reinforcement learning algorithms optimize the policy with respect to the expected gain. We propose an algorithm which enables to optimize an alternative objective: the probability that the gain is…
The mutual information (MI) between two random variables is an important correlation measure in data analysis. The Shannon entropy of a joint probability distribution is the variable part under fixed marginals. We aim to minimize and…
Maximal inequalities refer to bounds on expected values of the supremum of averages of random variables over a collection. They play a crucial role in the study of non-parametric and high-dimensional estimators, and especially in the study…
In this work, we consider a recently proposed entropy S (called varentropy) defined by a variational relationship dI=beta*(d<x>-<dx>) as a measure of uncertainty of random variable x. By definition, varentropy underlies a generalized…
In this paper, we present a maximum likelihood estimation approach to determine the value vector in transformer models. We model the sequence of value vectors, key vectors, and the query vector as a sequence of Gaussian distributions. The…
A variational principle is further developed for out of equilibrium dynamical systems by using the concept of maximum entropy. With this new formulation it is obtained a set of two first-order differential equations, revealing the same…
The well known maximum-entropy principle due to Jaynes, which states that given mean parameters, the maximum entropy distribution matching them is in an exponential family, has been very popular in machine learning due to its "Occam's…