Related papers: On Missing Mass Variance
The notion of uncertainty is of major importance in machine learning and constitutes a key element of machine learning methodology. In line with the statistical tradition, uncertainty has long been perceived as almost synonymous with…
Message importance measure (MIM) is applicable to characterize the importance of information in the scenario of big data, similar to entropy in information theory. In fact, MIM with a variable parameter can make an effect on the…
The nature of randomness in disordered packings of frictional and frictionless spheres is investigated using theory and simulations of identical spherical grains. The entropy of the packings is defined through the force and volume ensemble…
We assume that the forecast error follows a probability distribution which is symmetric and monotonically non-increasing on non-negative real numbers, and if there is a mismatch between observed and predicted value, then we suffer a loss.…
A model is proposed that describes the evolution of a mixed state of a quantum system for which gain and loss of energy or amplitude are present. Properties of the model are worked out in detail. In particular, invariant subspaces of the…
The defect of a function $f:M\rightarrow \mathbb{R}$ is defined as the difference between the measure of the positive and negative regions. In this paper, we begin the analysis of the distribution of defect of random Gaussian spherical…
Consider a measurement in which the current coming out of a mesoscopic sample is filtered around a given frequency, amplified, measured and squared. Then this process is repeated many times and the results are averaged. Often, two such…
Magnitude, obtained as a special case of Euler characteristic of enriched category, represents a sense of the size of metric spaces and is related to classical notions such as cardinality, dimension, and volume. While the studies have…
We relate the information entropy and the mass variance of any distribution in the regime of small fluctuations. We use a set of Monte Carlo simulations of different homogeneous and inhomogeneous distributions to verify the relation and…
Packing density is a permutation occurrence statistic which describes the maximal number of permutations of a given type that can occur in another permutation. In this article we focus on containment of sets of permutations. Although this…
Power-law distributions are typical macroscopic features occurring in almost all complex systems observable in nature. As a result, researchers in quantitative analyses must often generate random synthetic variates obeying power-law…
The pervasive presence in space of a flux of high-speed, electrically uncharged dark matter particles is examined here for potential consequences. Dark matter interactions with ordinary matter are considered, and a model of the dark matter…
We consider a multinomial distribution, where the number of cells increases and the cell-probabilities decreases as the number of observations grows. The probabilities of large deviations of statistics, which has form of a sum of Borel…
We present a mathematical method to statistically decouple the effects of unknown inclination angles on the mass distribution of exoplanets that have been discovered using radial-velocity techniques. The method is based on the distribution…
Majorisation, also called rearrangement inequalities, yields a type of stochastic ordering in which two or more distributions can be compared. In this paper we argue that majorisation is a good candidate as a theory for uncertainty. We…
Importance sampling is a popular variance reduction method for Monte Carlo estimation, where a notorious question is how to design good proposal distributions. While in most cases optimal (zero-variance) estimators are theoretically…
From the sampling of data to the initialisation of parameters, randomness is ubiquitous in modern Machine Learning practice. Understanding the statistical fluctuations engendered by the different sources of randomness in prediction is…
A researcher is interested in what sample size is needed to get the required significance of the same test, assuming exactly the same situation that was in the study with the non-significant result. We propose a simple solution to the…
To overcome the drawbacks of Shannon's entropy, the concept of cumulative residual and past entropy has been proposed in the information theoretic literature. Furthermore, the Shannon entropy has been generalized in a number of different…
Let A be finite set equipped with a probability distribution P, and let M be a "mass" function on A. A characterization is given for the most efficient way in which A^n can be covered using spheres of a fixed radius. A covering is a subset…