English
Related papers

Related papers: A deterministic and computable Bernstein-von Mises…

200 papers

We explore the computational content of Kronecker's lemma via the proof-theoretic perspective of proof mining and utilise the resulting finitary variant of this fundamental result to provide new rates for the Strong Law of Large Numbers for…

Logic · Mathematics 2024-11-14 Morenikeji Neri

We study the problem of maximum likelihood estimation of densities that are log-concave and lie in the graphical model corresponding to a given undirected graph $G$. We show that the maximum likelihood estimate (MLE) is the product of the…

Statistics Theory · Mathematics 2025-12-02 Kaie Kubjas , Olga Kuznetsova , Elina Robeva , Pardis Semnani , Luca Sodomaco

Estimating the Kullback--Leibler (KL) divergence between language models has many applications, e.g., reinforcement learning from human feedback (RLHF), interpretability, and knowledge distillation. However, computing the exact KL…

Computation and Language · Computer Science 2025-10-28 Afra Amini , Tim Vieira , Ryan Cotterell

We give the first mathematically rigorous justification of the Local Density Approximation in Density Functional Theory. We provide a quantitative estimate on the difference between the grand-canonical Levy-Lieb energy of a given density…

Mathematical Physics · Physics 2019-11-13 Mathieu Lewin , Elliott H. Lieb , Robert Seiringer

We interpret likelihood-based test functions from a geometric perspective where the Kullback-Leibler (KL) divergence is adopted to quantify the distance from a distribution to another. Such a test function can be seen as a sub-Gaussian…

Information Theory · Computer Science 2021-01-05 Yan Wang

The purpose of this paper is twofold. On a technical side, we propose an extension of the Hausdorff distance from metric spaces to spaces equipped with asymmetric distance measures. Specifically, we focus on the family of Bregman…

Machine Learning · Computer Science 2025-04-11 Tuyen Pham , Hana Dal Poz Kouřimská , Hubert Wagner

We present non-asymptotic two-sided bounds to the log-marginal likelihood in Bayesian inference. The classical Laplace approximation is recovered as the leading term. Our derivation permits model misspecification and allows the parameter…

Statistics Theory · Mathematics 2020-06-23 Anirban Bhattacharya , Debdeep Pati

We extend several recent results providing symmetry-based guarantees for variational inference (VI) with location-scale families. VI approximates a target density $p$ by the best match $q^*$ in a family $Q$ of tractable distributions that…

Machine Learning · Statistics 2025-12-11 Charles C. Margossian , Lawrence K. Saul

We study predictive density estimation under Kullback-Leibler loss in $\ell_0$-sparse Gaussian sequence models. We propose proper Bayes predictive density estimates and establish asymptotic minimaxity in sparse models. A surprise is the…

Statistics Theory · Mathematics 2017-08-01 Gourab Mukherjee , Iain M. Johnstone

We study nonparametric maximum likelihood estimation of a log-concave density function $f_0$ which is known to satisfy further constraints, where either (a) the mode $m$ of $f_0$ is known, or (b) $f_0$ is known to be symmetric about a fixed…

Statistics Theory · Mathematics 2019-05-15 Charles R. Doss , Jon A. Wellner

In this paper, we develop a general approach to proving global and local uniform limit theorems for the Horvitz-Thompson empirical process arising from complex sampling designs. Global theorems such as Glivenko-Cantelli and Donsker…

Statistics Theory · Mathematics 2019-05-31 Qiyang Han , Jon A. Wellner

The population $\mathrm{KL}_{\inf}$ is a fundamental quantity that appears in lower bounds for (asymptotically) optimal regret of pure-exploration stochastic bandit algorithms, and optimal stopping time of sequential tests. Motivated by…

Statistics Theory · Mathematics 2026-02-06 Ashwin Ram , Aaditya Ramdas

The LogSumExp function, dual to the Kullback-Leibler (KL) divergence, plays a central role in many important optimization problems, including entropy-regularized optimal transport (OT) and distributionally robust optimization (DRO). In…

Optimization and Control · Mathematics 2026-02-04 Egor Gladin , Alexey Kroshnin , Jia-Jie Zhu , Pavel Dvurechensky

Motivated by the computation of the non-parametric maximum likelihood estimator (NPMLE) and the Bayesian posterior in statistics, this paper explores the problem of convex optimization over the space of all probability distributions. We…

Statistics Theory · Mathematics 2023-11-03 Rentian Yao , Linjun Huang , Yun Yang

Variational Bayes (VB) provides a computationally efficient alternative to Markov Chain Monte Carlo, especially for high-dimensional and large-scale inference. However, existing theory on VB primarily focuses on fixed-dimensional settings…

Statistics Theory · Mathematics 2025-08-05 Jiawei Yan , Peirong Xu , Tao Wang

Let $X_1,\dots,X_n$ be i.i.d. log-concave random vectors in $\mathbb R^d$ with mean 0 and covariance matrix $\Sigma$. We study the problem of quantifying the normal approximation error for $W=n^{-1/2}\sum_{i=1}^nX_i$ with explicit…

Probability · Mathematics 2023-05-30 Xiao Fang , Yuta Koike

Consider the Gaussian sequence model under the additional assumption that a fixed fraction of the means is known. We study the problem of variance estimation from a frequentist Bayesian perspective. The maximum likelihood estimator (MLE)…

Statistics Theory · Mathematics 2019-12-19 Gianluca Finocchio , Johannes Schmidt-Hieber

In this paper, we study the statistical and geometrical properties of the Kullback-Leibler divergence with kernel covariance operators (KKL) introduced by Bach [2022]. Unlike the classical Kullback-Leibler (KL) divergence that involves…

Machine Learning · Statistics 2025-03-12 Clémentine Chazal , Anna Korba , Francis Bach

Classical (or ``global'') Bernstein theory establishes sharp control on entire functions of exponential type that are bounded and real-valued on the real axis. We localize some of this theory to rectangular regions $\{ x+iy: x \in I, 0 \leq…

Classical Analysis and ODEs · Mathematics 2026-04-23 Terence Tao

We discuss Bayesian methods for learning Bayesian networks when data sets are incomplete. In particular, we examine asymptotic approximations for the marginal likelihood of incomplete data given a Bayesian network. We consider the Laplace…

Machine Learning · Computer Science 2015-05-19 David Maxwell Chickering , David Heckerman