Related papers: On the Properties of Kullback-Leibler Divergence B…
The Kullback-Leibler divergence offers an information-theoretic basis for measuring the difference between two given distributions. Its quantum analog, however, fails to play a corresponding role for comparing two density matrices, if the…
Log-likelihood vectors define a common space for comparing language models as probability distributions, enabling unified comparisons across heterogeneous settings. We extend this framework to training checkpoints and intermediate layers,…
We derive tight non-asymptotic bounds for the Kolmogorov distance between the probabilities of two Gaussian elements to hit a ball in a Hilbert space. The key property of these bounds is that they are dimension-free and depend on the…
We report a closed-form expression for the Kullback-Leibler divergence between Cauchy distributions which involves the calculation of a novel definite integral. The formula shows that the Kullback-Leibler divergence between Cauchy densities…
Extreme Value Theory plays an important role to provide approximation results for the extremes of a sequence of independent random variables when their distribution is unknown. An important one is given by the {generalised Pareto…
Supplement 1 to GUM (GUM-S1) recommends the use of maximum entropy principle (MaxEnt) in determining the probability distribution of a quantity having specified properties, e.g., specified central moments. When we only know the mean value…
This paper proposes and studies new quantum version of $f$-divergences, a class of convex functionals of a pair of probability distributions including Kullback-Leibler divergence, Rnyi-type relative entropy and so on. There are several…
This paper shows that large nonparametric classes of conditional multivariate densities can be approximated in the Kullback--Leibler distance by different specifications of finite mixtures of normal regressions in which normal means and…
Gaussian mixture models (GMM) are powerful parametric tools with many applications in machine learning and computer vision. Expectation maximization (EM) is the most popular algorithm for estimating the GMM parameters. However, EM…
In this study, simultaneous predictive distributions for independent Poisson observables were considered and the performance of predictive distributions was evaluated using the Kullback-Leibler (K-L) loss. This study proposes a class of…
Small-scale intermittency is studied as the deviation of the probability distributions of pseudodissipation, dissipation and enstrophy in turbulence from those of a Gaussian random velocity field. This deviation is quantified using…
We develop a fully non-parametric, easy-to-use, and powerful test for the missing completely at random (MCAR) assumption on the missingness mechanism of a dataset. The test compares distributions of different missing patterns on random…
Motivated by applications in deep learning, where the global Lipschitz continuity condition is often not satisfied, we examine the problem of sampling from distributions with super-linearly growing log-gradients. We propose a novel tamed…
The maximum likelihood method is the best-known method for estimating the probabilities behind the data. However, the conventional method obtains the probability model closest to the empirical distribution, resulting in overfitting. Then…
We study density estimation in Kullback-Leibler divergence: given an i.i.d. sample from an unknown density $p^\star$, the goal is to construct an estimator $\widehat{p}$ such that $\mathrm{KL}(p^\star,\widehat{p})$ is small with high…
This paper introduces two new robust methods for estimation of parameters in a given parametric family. The first method is that of `minimum weighted L2', effectively minimising an estimate of the integrated (and possibly weighted) squared…
We prove an almost constant lower bound of the isoperimetric coefficient in the KLS conjecture. The lower bound has the dimension dependency $d^{-o_d(1)}$. When the dimension is large enough, our lower bound is tighter than the previous…
Knowing if a model will generalize to data 'in the wild' is crucial for safe deployment. To this end, we study model disagreement notions that consider the full predictive distribution - specifically disagreement based on Hellinger…
How much one has learned from an experiment is quantifiable by the information gain, also known as the Kullback-Leibler divergence. The narrowing of the posterior parameter distribution $P(\theta|D)$ compared with the prior parameter…
In this paper, we propose some estimators for the parameters of a statistical model based on Kullback-Leibler divergence of the survival function in continuous setting. We prove that the proposed estimators are subclass of "generalized…