English
Related papers

Related papers: Algorithmic randomness and the weak merging of com…

200 papers

The purpose of this paper is twofold. On a technical side, we propose an extension of the Hausdorff distance from metric spaces to spaces equipped with asymmetric distance measures. Specifically, we focus on the family of Bregman…

Machine Learning · Computer Science 2025-04-11 Tuyen Pham , Hana Dal Poz Kouřimská , Hubert Wagner

A loss function measures the discrepancy between the true values and their estimated fits, for a given instance of data. In classification problems, a loss function is said to be proper if a minimizer of the expected loss is the true…

Information Theory · Computer Science 2020-01-03 Amichai Painsky , Gregory W. Wornell

In optimization, the natural gradient method is well-known for likelihood maximization. The method uses the Kullback-Leibler divergence, corresponding infinitesimally to the Fisher-Rao metric, which is pulled back to the parameter space of…

Machine Learning · Statistics 2019-02-26 Anton Mallasto , Tom Dela Haije , Aasa Feragen

$f$-divergences are a general class of divergences between probability measures which include as special cases many commonly used divergences in probability, mathematical statistics and information theory such as Kullback-Leibler…

Statistics Theory · Mathematics 2013-10-16 Adityanand Guntuboyina , Sujayam Saha , Geoffrey Schiebinger

We show that the log-likelihood of several probabilistic graphical models is Lipschitz continuous with respect to the lp-norm of the parameters. We discuss several implications of Lipschitz parametrization. We present an upper bound of the…

Machine Learning · Computer Science 2018-11-16 Jean Honorio

Optimum designs for parameter estimation in generalized regression models are standardly based on the Fisher information matrix (cf. Atkinson et al (2014) for a recent exposition). The corresponding optimality criteria are related to the…

Statistics Theory · Mathematics 2015-07-28 Katarína Burclová , Andrej Pázman

The likelihood function is a fundamental component in Bayesian statistics. However, evaluating the likelihood of an observation is computationally intractable in many applications. In this paper, we propose a non-parametric approximation of…

Machine Learning · Computer Science 2019-10-24 Viet Anh Nguyen , Soroosh Shafieezadeh-Abadeh , Man-Chung Yue , Daniel Kuhn , Wolfram Wiesemann

Estimating Kullback Leibler (KL) divergence from samples of two distributions is essential in many machine learning problems. Variational methods using neural network discriminator have been proposed to achieve this task in a scalable…

Machine Learning · Computer Science 2021-10-01 Sandesh Ghimire , Aria Masoomi , Jennifer Dy

This paper introduces two new robust methods for estimation of parameters in a given parametric family. The first method is that of `minimum weighted L2', effectively minimising an estimate of the integrated (and possibly weighted) squared…

Methodology · Statistics 2026-02-23 Nils Lid Hjort

In this paper, we derive a useful lower bound for the Kullback-Leibler divergence (KL-divergence) based on the Hammersley-Chapman-Robbins bound (HCRB). The HCRB states that the variance of an estimator is bounded from below by the…

Statistics Theory · Mathematics 2019-11-05 Tomohiro Nishiyama

Good robust estimators can be tuned to combine a high breakdown point and a specified asymptotic efficiency at a central model. This happens in regression with MM- and tau-estimators among others. However, the finite-sample efficiency of…

Statistics Theory · Mathematics 2013-11-21 Ricardo Maronna , Víctor Yohai

We consider the problem of constructing a least conservative estimator of the expected value $\mu$ of a non-negative heavy-tailed random variable. We require that the probability of overestimating the expected value $\mu$ is kept…

Optimization and Control · Mathematics 2026-04-21 Bart P. G. van Parys , Bert Zwart

Estimating the Kullback--Leibler (KL) divergence between language models has many applications, e.g., reinforcement learning from human feedback (RLHF), interpretability, and knowledge distillation. However, computing the exact KL…

Computation and Language · Computer Science 2025-10-28 Afra Amini , Tim Vieira , Ryan Cotterell

Discrete normal distributions are defined as the distributions with prescribed means and covariance matrices which maximize entropy on the integer lattice support. The set of discrete normal distributions form an exponential family with…

Information Theory · Computer Science 2022-01-25 Frank Nielsen

We study generalizations of Demuth's Theorem, which states that the image of a Martin-L\"of random real under a tt-reduction is either computable or Turing equivalent to a Martin-L\"of random real. We show that Demuth's Theorem holds for…

Logic · Mathematics 2011-10-27 Laurent Bienvenu , Christopher Porter

We study Martin-L\"{o}f random (ML-random) points on computable probability measures on sample and parameter spaces (Bayes models). We consider variants of conditional randomness defined by ML-randomness on Bayes models and those of…

Information Theory · Computer Science 2023-04-24 Hayato Takahashi

The Kullback-Leibler divergence, the Kullback-Leibler variation, and the Bernstein "norm" are used to quantify discrepancies among probability distributions in likelihood models such as nonparametric maximum likelihood and nonparametric…

Statistics Theory · Mathematics 2026-01-27 Tetsuya Kaji

This paper shows that large nonparametric classes of conditional multivariate densities can be approximated in the Kullback--Leibler distance by different specifications of finite mixtures of normal regressions in which normal means and…

Statistics Theory · Mathematics 2010-10-05 Andriy Norets

In this work we introduce a family of transformations, named \textit{divergence transformations}, interpolating between any pair of probability density functions sharing the same support. We prove the remarkable property that the whole…

Mathematical Physics · Physics 2025-12-15 Razvan Gabriel Iagar , David Puertas-Centeno , Elio V. Toranzo

Algorithmic randomness theory starts with a notion of an individual random object. To be reasonable, this notion should have some natural properties; in particular, an object should be random with respect to image distribution if and only…

Logic · Mathematics 2016-07-15 Laurent Bienvenu , Mathieu Hoyrup , Alexander Shen