English
Related papers

Related papers: Optimal Kullback-Leibler Aggregation in Mixture De…

200 papers

The Kullback-Leibler (KL) divergence plays a central role in probabilistic machine learning, where it commonly serves as the canonical loss function. Optimization in such settings is often performed over the probability simplex, where the…

Machine Learning · Computer Science 2025-07-31 Adwait Datar , Nihat Ay

Detection of sparse signals arises in a wide range of modern scientific studies. The focus so far has been mainly on Gaussian mixture models. In this paper, we consider the detection problem under a general sparse mixture model and obtain…

Information Theory · Computer Science 2012-11-13 T. Tony Cai , Yihong Wu

Introduced by Kiefer and Wolfowitz \cite{KW56}, the nonparametric maximum likelihood estimator (NPMLE) is a widely used methodology for learning mixture odels and empirical Bayes estimation. Sidestepping the non-convexity in mixture…

Statistics Theory · Mathematics 2020-09-08 Yury Polyanskiy , Yihong Wu

We consider the problem of sampling from a probability distribution $\pi$ which admits a density w.r.t. a dominating measure. It is well known that this can be written as an optimisation problem over the space of probability distributions…

Methodology · Statistics 2026-05-06 Francesca Romana Crucinio

Frequentist conditions for asymptotic suitability of Bayesian procedures focus on lower bounds for prior mass in Kullback-Leibler neighbourhoods of the data distribution. The goal of this paper is to investigate the flexibility in criteria…

Statistics Theory · Mathematics 2018-03-19 B. J. K. Kleijn , Y. Y. Zhao

Diffusion models are a new class of generative models that revolve around the estimation of the score function associated with a stochastic differential equation. Subsequent to its acquisition, the approximated score function is then…

Statistics Theory · Mathematics 2024-09-13 Giovanni Conforti , Alain Durmus , Marta Gentiloni Silveri

In contingency table analysis, sparse data is frequently encountered for even modest numbers of variables, resulting in non-existence of maximum likelihood estimates. A common solution is to obtain regularized estimates of the parameters of…

Methodology · Statistics 2015-11-04 James E. Johndrow , Anirban Bhattacharya

We emphasize that it is possible to improve the principle of unbiased risk estimation for model selection by addressing excess risk deviations in the design of penalization procedures. Indeed, we propose a modification of Akaike's…

Statistics Theory · Mathematics 2018-07-23 Adrien Saumard , Fabien Navarro

Data collection is a critical step in statistical inference and data science, and the goal of statistical experimental design (ED) is to find the data collection setup that can provide most information for the inference. In this work we…

Computation · Statistics 2020-07-01 Ziqiao Ao , Jinglai Li

The problem of estimation of the distribution parameters on the sample when the part of these parameters are discrete (e.g. integer) is considered. We prove that the rate of convergence of MLE estimates under the natural conditions on the…

Statistics Theory · Mathematics 2014-02-27 E. Ostrovsky , L. Sirota , A. Zeldin

In optimization, the natural gradient method is well-known for likelihood maximization. The method uses the Kullback-Leibler divergence, corresponding infinitesimally to the Fisher-Rao metric, which is pulled back to the parameter space of…

Machine Learning · Statistics 2019-02-26 Anton Mallasto , Tom Dela Haije , Aasa Feragen

In this paper we study the computation of the nonparametric maximum likelihood estimator (NPMLE) in multivariate mixture models. Our first approach discretizes this infinite dimensional convex optimization problem by fixing the support…

Methodology · Statistics 2024-02-20 Yangjing Zhang , Ying Cui , Bodhisattva Sen , Kim-Chuan Toh

We study sparsity-regularized maximum likelihood estimation for the drift parameter of high-dimensional non-stationary Ornstein--Uhlenbeck processes given repeated measurements of i.i.d. paths. In particular, we show that Lasso and Slope…

Statistics Theory · Mathematics 2025-10-29 Shogo Nakakita

In this paper we study the frequentist convergence rate for the Latent Dirichlet Allocation (Blei et al., 2003) topic models. We show that the maximum likelihood estimator converges to one of the finitely many equivalent parameters in…

Machine Learning · Statistics 2019-01-21 Yining Wang

We construct optimal low-rank approximations for the Gaussian posterior distribution in linear Gaussian inverse problems with possibly infinite-dimensional separable Hilbert parameter spaces and finite-dimensional data spaces. We first…

Statistics Theory · Mathematics 2026-04-09 Giuseppe Carere , Han Cheng Lie

This paper proposes a new family of lower and upper bounds on the minimum mean squared error (MMSE). The key idea is to minimize/maximize the MMSE subject to the constraint that the joint distribution of the input-output statistics lies in…

Information Theory · Computer Science 2020-06-09 Michael Fauß , Alex Dysto , H. Vincent Poor

We propose a new estimation procedure of the conditional density for independent and identically distributed data. Our procedure aims at using the data to select a function among arbitrary (at most countable) collections of candidates. By…

Statistics Theory · Mathematics 2016-10-26 Mathieu Sart

For a regression model, we consider the risk of the maximum likelihood estimator with respect to $\alpha$-divergence, which includes the special cases of Kullback-Leibler divergence, Hellinger distance and $\chi^2$ divergence. The…

Statistics Theory · Mathematics 2017-09-12 Yo Sheena

The performance of machine learning classification algorithms are evaluated by estimating metrics, often from the confusion matrix, using training data and cross-validation. However, these do not prove that the best possible performance has…

Machine Learning · Statistics 2024-03-05 L. Crow , S. J. Watts

The standard paradigm of neural language generation adopts maximum likelihood estimation (MLE) as the optimizing method. From a distributional view, MLE in fact minimizes the Kullback-Leibler divergence (KLD) between the distribution of the…

Computation and Language · Computer Science 2023-02-28 Haozhe Ji , Pei Ke , Zhipeng Hu , Rongsheng Zhang , Minlie Huang