Related papers: On the R\'{e}nyi Cross-Entropy
Shannon entropy is the shortest average codeword length a lossless compressor can achieve by encoding i.i.d. symbols. However, there are cases in which the objective is to minimize the \textit{exponential} average codeword length, i.e. when…
Estimating the entropy of a discrete random variable is a fundamental problem in information theory and related fields. This problem has many applications in various domains, including machine learning, statistics and data compression. Over…
The R{\'e}nyi entropy is one of the important information measures that generalizes Shannon's entropy. The quantum R{\'e}nyi entropy has a fundamental role in quantum information theory, therefore, bounding this quantity is of vital…
Recently, substantial research efforts in Deep Metric Learning (DML) focused on designing complex pairwise-distance losses, which require convoluted schemes to ease optimization, such as sample mining or pair weighting. The standard…
Entanglement criteria for an $n$-partite quantum system with continuous variables are formulated in terms of R\'{e}nyi entropies. R\'{e}nyi entropies are widely used as a good information measure due to many nice properties. Derived…
We propose "collision cross-entropy" as a robust alternative to Shannon's cross-entropy (CE) loss when class labels are represented by soft categorical distributions y. In general, soft labels can naturally represent ambiguous targets in…
We describe a method to estimate R\'enyi entanglement entropy of a spin system, which is based on the replica trick and generative neural networks with explicit probability estimation. It can be extended to any spin system or lattice field…
R\'enyi transfer entropy (RTE) is a generalization of classical transfer entropy that replaces Shannon's entropy with R\'enyi's information measure. This, in turn, introduces a new tunable parameter $\alpha$, which accounts for sensitivity…
Shannon entropy for discrete distributions is a fundamental and widely used concept, but its continuous analogue, known as differential entropy, lacks essential properties such as positivity and compatibility with the discrete case. In this…
We show that the R\'enyi entropies of single particle, extended wave functions for disordered systems contain information about the multifractal spectrum. It is shown for moments of the R\'enyi entropy, $S_{n}$, where $|n|<1$, it is…
We propose a permutation-invariant loss function designed for the neural networks reconstructing a set of elements without considering the order within its vector representation. Unlike popular approaches for encoding and decoding a set,…
In this study an attempt has been made to propose a way to develop new distribution. For this purpose, we need only idea about distribution function. Some important statistical properties of the new distribution like moments, cumulants,…
In this paper, we focus on the separability of classes with the cross-entropy loss function for classification problems by theoretically analyzing the intra-class distance and inter-class distance (i.e. the distance between any two points…
Nearly all practical neural models for classification are trained using cross-entropy loss. Yet this ubiquitous choice is supported by little theoretical or empirical evidence. Recent work (Hui & Belkin, 2020) suggests that training using…
We prove that all R\'enyi entanglement entropies of spin-chains described by generic (gapped), translational invariant matrix product states (MPS) are extensive for disconnected sub-systems: All R\'enyi entanglement entropy densities of the…
We derive a family of entanglement criteria for continuous variable systems based on the R\'enyi entropy of complementary distributions. We show that these entanglement witnesses can be more sensitive than those based on second-order…
Shannon Entropy is the preeminent tool for measuring the level of uncertainty (and conversely, information content) in a random variable. In the field of communications, entropy can be used to express the information content of given…
Feature selection, in the context of machine learning, is the process of separating the highly predictive feature from those that might be irrelevant or redundant. Information theory has been recognized as a useful concept for this task, as…
There is no such thing as a perfect dataset. In some datasets, deep neural networks discover underlying heuristics that allow them to take shortcuts in the learning process, resulting in poor generalization capability. Instead of using…
One of the most useful tools for distinguishing between chaotic and stochastic time series is the so-called complexity-entropy causality plane. This diagram involves two complexity measures: the Shannon entropy and the statistical…