English
Related papers

Related papers: Existence of Direct Density Ratio Estimators

200 papers

We propose new model selection criteria based on generalized ridge estimators dominating the maximum likelihood estimator under the squared risk and the Kullback-Leibler risk in multivariate linear regression. Our model selection criteria…

Statistics Theory · Mathematics 2016-04-08 Yuichi Mori , Taiji Suzuki

We address the problem of detecting a change in the distribution of a high-dimensional multivariate normal time series. Assuming that the post-change parameters are unknown and estimated using a window of historical data, we extend the…

Signal Processing · Electrical Eng. & Systems 2025-02-12 Robert Malinas , Dogyoon Song , Benjamin D. Robinson , Alfred O. Hero

Information divergences allow one to assess how close two distributions are from each other. Among the large panel of available measures, a special attention has been paid to convex $\varphi$-divergences, such as Kullback-Leibler,…

Information Theory · Computer Science 2019-04-09 Mireille El Gheche , Giovanni Chierchia , Jean-Christophe Pesquet

In this work we introduce a family of transformations, named \textit{divergence transformations}, interpolating between any pair of probability density functions sharing the same support. We prove the remarkable property that the whole…

Mathematical Physics · Physics 2025-12-15 Razvan Gabriel Iagar , David Puertas-Centeno , Elio V. Toranzo

Bayesian inference with empirical likelihood faces a challenge as the posterior domain is a proper subset of the original parameter space due to the convex hull constraint. We propose a regularized exponentially tilted empirical likelihood…

Methodology · Statistics 2026-04-23 Eunseop Kim , Steven N. MacEachern , Mario Peruggia

In this paper, we propose a theoretical analysis of the algorithm ISDE, introduced in previous work. From a dataset, ISDE learns a density written as a product of marginal density estimators over a partition of the features. We show that…

Statistics Theory · Mathematics 2022-05-09 Louis Pujol

We study two adaptive importance sampling schemes for estimating the probability of a rare event in the high-dimensional regime $d \to \infty$ with $d$ the dimension. The first scheme is the prominent cross-entropy (CE) method, and the…

Statistics Theory · Mathematics 2025-03-26 Jason Beh , Yonatan Shadmi , Florian Simatos

Bayesian learning has been recently considered as an effective means of accounting for uncertainty in trained deep network parameters. This is of crucial importance when dealing with small or sparse training datasets. On the other hand,…

Machine Learning · Computer Science 2018-02-13 Harris Partaourides , Sotirios Chatzis

This paper applies the recently axiomatized Optimum Information Principle (minimize the Kullback-Leibler information subject to all relevant information) to nonparametric density estimation, which provides a theoretical foundation as well…

Statistics Theory · Mathematics 2011-03-28 Alexis Akira Toda

Deep neural networks (DNNs) exhibit an exceptional capacity for generalization in practical applications. This work aims to capture the effect and benefits of depth for supervised learning via information-theoretic generalization bounds. We…

Machine Learning · Computer Science 2025-05-09 Haiyun He , Ziv Goldfeld

Several scalable sample-based methods to compute the Kullback Leibler (KL) divergence between two distributions have been proposed and applied in large-scale machine learning models. While they have been found to be unstable, the…

Machine Learning · Computer Science 2021-09-07 Sandesh Ghimire , Prashnna K Gyawali , Linwei Wang

Estimator selection has become a crucial issue in non parametric estimation. Two widely used methods are penalized empirical risk minimization (such as penalized log-likelihood estimation) or pairwise comparison (such as Lepski's method).…

Statistics Theory · Mathematics 2017-10-19 Claire Lacour , Pascal Massart , Vincent Rivoirard

Recently, continual learning has received a lot of attention. One of the significant problems is the occurrence of \emph{concept drift}, which consists of changing probabilistic characteristics of the incoming data. In the case of the…

Machine Learning · Computer Science 2022-10-11 Sebastián Basterrech , Michal Woźniak

Let $V_* : \mathbb{R}^d \to \mathbb{R}$ be some (possibly non-convex) potential function, and consider the probability measure $\pi \propto e^{-V_*}$. When $\pi$ exhibits multiple modes, it is known that sampling techniques based on…

Optimization and Control · Mathematics 2023-02-24 Carles Domingo-Enrich , Aram-Alexandre Pooladian

Designing experiments that systematically gather data from complex physical systems is central to accelerating scientific discovery. While Bayesian experimental design (BED) provides a principled, information-based framework that integrates…

Machine Learning · Computer Science 2026-01-26 Huchen Yang , Xinghao Dong , Jin-Long Wu

The problem of estimating the Kullback-Leibler divergence $D(P\|Q)$ between two unknown distributions $P$ and $Q$ is studied, under the assumption that the alphabet size $k$ of the distributions can scale to infinity. The estimation is…

Information Theory · Computer Science 2018-02-22 Yuheng Bu , Shaofeng Zou , Yingbin Liang , Venugopal V. Veeravalli

In this paper, we consider a distributionally robust optimization (DRO) model in which the ambiguity set is defined as the set of distributions whose Kullback-Leibler (KL) divergence to an empirical distribution is bounded. Utilizing the…

Optimization and Control · Mathematics 2024-11-12 Burak Kocuk

Recent Reinforcement Learning (RL) algorithms making use of Kullback-Leibler (KL) regularization as a core component have shown outstanding performance. Yet, only little is understood theoretically about why KL regularization helps, so far.…

Machine Learning · Computer Science 2021-01-07 Nino Vieillard , Tadashi Kozuno , Bruno Scherrer , Olivier Pietquin , Rémi Munos , Matthieu Geist

We study the problem of estimating a distribution over a finite alphabet from an i.i.d. sample, with accuracy measured in relative entropy (Kullback-Leibler divergence). While optimal bounds on the expected risk are known, high-probability…

Statistics Theory · Mathematics 2026-02-27 Jaouad Mourtada

Estimating the Kullback--Leibler (KL) divergence between language models has many applications, e.g., reinforcement learning from human feedback (RLHF), interpretability, and knowledge distillation. However, computing the exact KL…

Computation and Language · Computer Science 2025-10-28 Afra Amini , Tim Vieira , Ryan Cotterell