Related papers: New Upper bounds for KL-divergence Based on Integr…
We present a statistical mechanical framework based on the Kullback-Leibler divergence (KLD) to analyze the relativistic limits of decoding time-encoded information from a moving source. By modeling the symbol durations as…
We establish some new non-asymptotical lower bounds for deviation of regular unbiased estimation of unknown parameter from its true value in different norms, alike the classical Rao-Kramer's inequality. We show that if the new norm is…
Estimating the Kullback--Leibler (KL) divergence between language models has many applications, e.g., reinforcement learning from human feedback (RLHF), interpretability, and knowledge distillation. However, computing the exact KL…
This work explores connections between the quantum relative entropy of two faithful states $\rho,\sigma$ (i.e. full-rank density matrices) and the Kullback-Leibler divergences of classical measures $\mu,\nu$. Here, $\mu$ and $\nu$ are…
We provide estimates of the rate of strong approximation and bounds for probabilities of moderate deviations in the CLT for the $L_1$-norm of the kernel density estimator without any assumptions on the density and assuming that the kernel…
The necessary information to distinguish a local inhomogeneous mass density field from its spatial average on a compact domain of the universe can be measured by relative information entropy. The Kullback-Leibler (KL) formula arises very…
This paper explores the connections between tempering (for Sequential Monte Carlo; SMC) and entropic mirror descent to sample from a target probability distribution whose unnormalized density is known. We establish that tempering SMC…
This paper proposes a new family of lower and upper bounds on the minimum mean squared error (MMSE). The key idea is to minimize/maximize the MMSE subject to the constraint that the joint distribution of the input-output statistics lies in…
Knowing if a model will generalize to data 'in the wild' is crucial for safe deployment. To this end, we study model disagreement notions that consider the full predictive distribution - specifically disagreement based on Hellinger…
The Katz-Sarnak density conjecture states that, as the analytic conductor $R \to \infty$, the distribution of the normalized low-lying zeros (those near the central point $s = 1/2$) converges to the scaling limits of eigenvalues clustered…
In this paper, we verify the $L^2$-boundedness for the jump functions and variations of Calder\'on-Zygmund singular integral operators with the underlying kernels satisfying \begin{align*}\int_{\varepsilon\leq |x-y|\leq N}…
A generalized Kullback-Leibler relative entropy is introduced starting with the symmetric Jackson derivative of the generalized overlap between two probability distributions. The generalization retains much of the structure possessed by the…
Consider the Langevin diffusion process $\mathrm{d} X_t = \nabla \log p_t(X_t) + \sqrt{2}\mathrm{d} W_t$ guided by the time-dependent probability density $p_t(x)$. Let $q_t$ be the density of $X_t$. Recently, in order to analyze convergence…
Trajectory Inference (TI) seeks to recover latent dynamical processes from snapshot data, where only independent samples from time-indexed marginals are observed. In applications such as single-cell genomics, destructive measurements make…
Universal hypothesis testing refers to the problem of deciding whether samples come from a nominal distribution or an unknown distribution that is different from the nominal distribution. Hoeffding's test, whose test statistic is equivalent…
$f$-divergences, which quantify discrepancy between probability distributions, are ubiquitous in information theory, machine learning, and statistics. While there are numerous methods for estimating $f$-divergences from data, a limit…
We study the problem of model selection type aggregation with respect to the Kullback-Leibler divergence for various probabilistic models. Rather than considering a convex combination of the initial estimators $f_1, \ldots, f_N$, our…
The paper addresses a problem of sampling discretization of integral norms of elements of finite-dimensional subspaces satisfying some conditions. We prove sampling discretization results under a standard assumption formulated in terms of…
Motivated by applications in deep learning, where the global Lipschitz continuity condition is often not satisfied, we examine the problem of sampling from distributions with super-linearly growing log-gradients. We propose a novel tamed…
For a map of the unit interval with an indifferent fixed point, we prove an upper bound for the variance of all observables of $n$ variables $K:[0,1]^n\to\R$ which are componentwise Lipschitz. The proof is based on coupling and decay of…