English
Related papers

Related papers: A deterministic and computable Bernstein-von Mises…

200 papers

Recently, a new distance has been introduced for the graphs of two point-to-set operators, one of which is maximally monotone. When both operators are the subdifferential of a proper lower semicontinuous convex function, this distance…

Functional Analysis · Mathematics 2020-09-29 Regina S. Burachik , Minh N. Dao , Scott B. Lindstrom

Maximum Likelihood Estimators (MLE) has many good properties. For example, the asymptotic variance of MLE solution attains equality of the asymptotic Cram{\'e}r-Rao lower bound (efficiency bound), which is the minimum possible variance for…

Machine Learning · Statistics 2019-11-05 Song Liu , Takafumi Kanamori , Wittawat Jitkrittum , Yu Chen

In contingency table analysis, sparse data is frequently encountered for even modest numbers of variables, resulting in non-existence of maximum likelihood estimates. A common solution is to obtain regularized estimates of the parameters of…

Methodology · Statistics 2015-11-04 James E. Johndrow , Anirban Bhattacharya

We study the gradient flow for a relaxed approximation to the Kullback-Leibler (KL) divergence between a moving source and a fixed target distribution. This approximation, termed the KALE (KL approximate lower-bound estimator), solves a…

Machine Learning · Statistics 2021-11-01 Pierre Glaser , Michael Arbel , Arthur Gretton

This paper investigates the behavior of statistical ensembles under iteration map induced by discrete integrable Hamiltonian systems in deterministic case and stochastic case, addressing the problem from two perspectives: the Law of Large…

Probability · Mathematics 2025-09-26 Xinyu Liu , Xinze Zhang , Yong Li

We consider a sparse linear regression model with unknown symmetric error under the high-dimensional setting. The true error distribution is assumed to belong to the locally $\beta$-H\"{o}lder class with an exponentially decreasing tail,…

Statistics Theory · Mathematics 2020-09-01 Kyoungjae Lee , Minwoo Chae , Lizhen Lin

Inferring and comparing complex, multivariable probability density functions is fundamental to problems in several fields, including probabilistic learning, network theory, and data analysis. Classification and prediction are the two faces…

Information Theory · Computer Science 2017-03-30 David J. Galas , T. Gregory Dewey , James Kunert-Graf , Nikita A. Sakhanenko

We study Bregman divergences in probability density space embedded with the $L^2$-Wasserstein metric. Several properties and dualities of transport Bregman divergences are provided. In particular, we derive the transport Kullback-Leibler…

Information Theory · Computer Science 2025-04-08 Wuchen Li

The Reifenberg theorem \cite{reif_orig} tells us that if a set $S\subseteq B_2\subseteq \mathbb R^n$ is uniformly close on all points and scales to a $k$-dimensional subspace, then $S$ is H\"older homeomorphic to a $k$-dimensional Euclidean…

Analysis of PDEs · Mathematics 2024-05-07 Nicholas Edelen , Aaron Naber , Daniele Valtorta

Universal hypothesis testing refers to the problem of deciding whether samples come from a nominal distribution or an unknown distribution that is different from the nominal distribution. Hoeffding's test, whose test statistic is equivalent…

Information Theory · Computer Science 2017-11-15 Pengfei Yang , Biao Chen

We consider Bayesian variable selection in sparse high-dimensional regression, where the number of covariates $p$ may be large relative to the samples size $n$, but at most a moderate number $q$ of covariates are active. Specifically, we…

Statistics Theory · Mathematics 2015-03-31 Rina Foygel Barber , Mathias Drton , Kean Ming Tan

When belief propagation (BP) converges, it does so to a stationary point of the Bethe free energy $F$, and is often strikingly accurate. However, it may converge only to a local optimum or may not converge at all. An algorithm was recently…

Machine Learning · Computer Science 2014-01-03 Adrian Weller , Tony Jebara

We study density estimation in Kullback-Leibler divergence: given an i.i.d. sample from an unknown density $p^\star$, the goal is to construct an estimator $\widehat{p}$ such that $\mathrm{KL}(p^\star,\widehat{p})$ is small with high…

Statistics Theory · Mathematics 2026-04-03 Spencer Compton , Gábor Lugosi , Jaouad Mourtada , Jian Qian , Nikita Zhivotovskiy

We give a new characterization of relative entropy, also known as the Kullback-Leibler divergence. We use a number of interesting categories related to probability theory. In particular, we consider a category FinStat where an object is a…

Information Theory · Computer Science 2017-08-22 John C. Baez , Tobias Fritz

Kullback-Leibler divergence (KL) regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise limits. Utilizing a unified information-geometric framework, we…

Optimization and Control · Mathematics 2026-02-03 Viktor Stein , Adwait Datar , Nihat Ay

Few methods in Bayesian non-parametric statistics/ machine learning have received as much attention as Bayesian Additive Regression Trees (BART). While BART is now routinely performed for prediction tasks, its theoretical properties began…

Statistics Theory · Mathematics 2019-05-10 Veronika Rockova

We consider the problem of predictive density estimation under Kullback-Leibler loss in a high-dimensional Gaussian model with exact sparsity constraints on the location parameters. We study the first order asymptotic minimax risk of Bayes…

Statistics Theory · Mathematics 2019-05-24 Ujan Gangopadhyay , Gourab Mukherjee

We establish effective convergence rates in the Doeblin-Lenstra law, describing the limiting distribution of approximation coefficients arising from continued fraction convergents of a typical real number. More generally, we prove…

Number Theory · Mathematics 2025-07-28 Gaurav Aggarwal , Anish Ghosh

Four problems related to information divergence measures defined on finite alphabets are considered. In three of the cases we consider, we illustrate a contrast which arises between the binary-alphabet and larger-alphabet settings. This is…

Information Theory · Computer Science 2016-11-17 Jiantao Jiao , Thomas Courtade , Albert No , Kartik Venkat , Tsachy Weissman

Bernstein's theorem (also called Hausdorff--Bernstein--Widder theorem) enables the integral representation of a completely monotonic function. We introduce a finite completely monotonic function, which is a completely monotonic function…

Numerical Analysis · Mathematics 2023-07-25 Yohei M. Koyama