Related papers: Trivergence of Probability Distributions, at glanc…
There are three classical divergence measures in the literature on information theory and statistics, namely, Jeffryes-Kullback-Leiber's J-divergence, Sibson-Burbea-Rao's Jensen-Shannon divegernce and Taneja's arithemtic-geometric mean…
Many machine learning tasks such as clustering, classification, and dataset search benefit from embedding data points in a space where distances reflect notions of relative similarity as perceived by humans. A common way to construct such…
How much one has learned from an experiment is quantifiable by the information gain, also known as the Kullback-Leibler divergence. The narrowing of the posterior parameter distribution $P(\theta|D)$ compared with the prior parameter…
$f$-divergences are a general class of divergences between probability measures which include as special cases many commonly used divergences in probability, mathematical statistics and information theory such as Kullback-Leibler…
In this paper we shall consider one parametric generalization of some non-symmetric divergence measures. The \textit{non-symmetric divergence measures} are such as: Kullback-Leibler \textit{relative information}, $\chi…
We establish quantitative comparisons between classical distances for probability distributions belonging to the class of convex probability measures. Distances include total variation distance, Wasserstein distance, Kullback-Leibler…
The modality is important topic for modelling. Using parametric models is an efficient way when real data set shows trimodality. In this paper we propose a new class of trimodal probability distributions, that is, probability distributions…
Statistical divergence is widely applied in multimedia processing, basically due to regularity and interpretable features displayed in data. However, in a broader range of data realm, these advantages may no longer be feasible, and…
In order to design haptic icons or build a haptic vocabulary, we require a set of easily distinguishable haptic signals to avoid perceptual ambiguity, which in turn requires a way to accurately estimate the perceptual (dis)similarity of…
We introduce a conceptually simple and effective method to quantify the similarity between relations in knowledge bases. Specifically, our approach is based on the divergence between the conditional probability distributions over entity…
The Kullback-Leibler (KL) divergence is frequently used in data science. For discrete distributions on large state spaces, approximations of probability vectors may result in a few small negative entries, rendering the KL divergence…
There are many applications that benefit from computing the exact divergence between 2 discrete probability measures, including machine learning. Unfortunately, in the absence of any assumptions on the structure or independencies within…
We propose a novel probabilistic approach to multilevel clustering problems based on composite transportation distance, which is a variant of transportation distance where the underlying metric is Kullback-Leibler divergence. Our method…
To what extent can we distinguish one probability distribution from another? Are there quantitative measures of distinguishability? The goal of this tutorial is to approach such questions by introducing the notion of the "distance" between…
Formalising the confrontation of opinions (models) to observations (data) is the task of Inferential Statistics. Information Theory provides us with a basic functional, the relative entropy (or Kullback-Leibler divergence), an asymmetrical…
Learning word representations has garnered greater attention in the recent past due to its diverse text applications. Word embeddings encapsulate the syntactic and semantic regularities of sentences. Modelling word embedding as multi-sense…
There are three classical divergence measures known in the literature on information theory and statistics. These are namely, Jeffryes-Kullback-Leiber \cite{jef} \cite{kul} \textit{J-divergence}. Sibson-Burbea-Rao \cite{sib} \cite{bur1,…
When sampling multi-modal probability distributions, correctly estimating the relative probability of each mode, even when the modes have been discovered and locally sampled, remains challenging. We test a simple reweighting scheme designed…
Here, we propose a new tool to estimate the complexity of a time series: the entropy of difference (ED). The method is based solely on the sign of the difference between neighboring values in a time series. This makes it possible to…
Given a probability distribution $\mu$ a set $\Lambda (\mu)$ of positive real numbers is introduced, so that $\Lambda (\mu)$ measures the "divisibility" of $\mu$. The basic properties of $\Lambda (\mu)$ are described and examples of…