English
Related papers

Related papers: Analysis of KNN Density Estimation

200 papers

We consider the fundamental problem of estimating a discrete distribution on a domain of size $K$ with high probability in Kullback-Leibler divergence. We provide upper and lower bounds on the minimax estimation rate, which show that the…

Machine Learning · Statistics 2026-02-23 Dirk van der Hoeven , Julia Olkhovskaia , Tim van Erven

We consider the task of estimating a conditional density using i.i.d. samples from a joint distribution, which is a fundamental problem with applications in both classification and uncertainty quantification for regression. For joint…

Statistics Theory · Mathematics 2023-06-16 Blair Bilodeau , Dylan J. Foster , Daniel M. Roy

We study the worst case error of kernel density estimates via subset approximation. A kernel density estimate of a distribution is the convolution of that distribution with a fixed kernel (e.g. Gaussian kernel). Given a subset (i.e. a point…

Computational Geometry · Computer Science 2012-04-05 Jeff M. Phillips

This paper presents minimax rates for density estimation when the data dimension $d$ is allowed to grow with the number of observations $n$ rather than remaining fixed as in previous analyses. We prove a non-asymptotic lower bound which…

Statistics Theory · Mathematics 2017-08-16 Daniel J. McDonald

We show that the cumulative distribution function corresponding to a kernel density estimator with optimal bandwidth lies outside any confidence interval, around the empirical distribution function, with probability tending to 1 as the…

Statistics Theory · Mathematics 2026-04-17 Nils Lid Hjort , Stephen G. Walker

We study nonparametric density estimation problems where error is measured in the Wasserstein distance, a metric on probability distributions popular in many areas of statistics and machine learning. We give the first minimax-optimal rates…

Statistics Theory · Mathematics 2020-04-30 Jonathan Niles-Weed , Quentin Berthet

Given a sample $\{X_i\}_{i=1}^n$ from $f_X$, we construct kernel density estimators for $f_Y$, the convolution of $f_X$ with a known error density $f_{\epsilon}$. This problem is known as density estimation with Berkson error and has…

Methodology · Statistics 2014-07-30 James P. Long , Noureddine El Karoui , John A. Rice

We study the problem of estimating a nonparametric probability density under a large family of losses called Besov IPMs, which include, for example, $\mathcal{L}^p$ distances, total variation distance, and generalizations of both…

Statistics Theory · Mathematics 2020-01-14 Ananya Uppal , Shashank Singh , Barnabás Póczos

Consider the problem of estimating the $\gamma$-level set $G^*_{\gamma}=\{x:f(x)\geq\gamma\}$ of an unknown $d$-dimensional density function $f$ based on $n$ independent observations $X_1,...,X_n$ from the density. This problem has been…

Statistics Theory · Mathematics 2009-08-26 Aarti Singh , Clayton Scott , Robert Nowak

In the k-nearest neighbor algorithm (k-NN), the determination of classes for test instances is usually performed via a majority vote system, which may ignore the similarities among data. In this research, the researcher proposes an approach…

Machine Learning · Computer Science 2019-06-13 Jasper Kyle Catapang

We examine the estimation of the Kullback-Leibler (KL) divergence and the use of the goodness-of-fit test for multivariate continuous distributions. Our starting point is the maximum entropy principle for Shannon entropy: among all…

Statistics Theory · Mathematics 2026-03-10 Mehmet Siddik Cadirci , Martin Singull

We consider an extension of $\epsilon$-entropy to a KL-divergence based complexity measure for randomized density estimation methods. Based on this extension, we develop a general information-theoretical inequality that measures the…

Statistics Theory · Mathematics 2007-06-13 Tong Zhang

The Bayes Error Rate (BER) is the fundamental limit on the achievable generalizable classification accuracy of any machine learning model due to inherent uncertainty within the data. BER estimators offer insight into the difficulty of any…

Machine Learning · Computer Science 2025-09-24 Lesley Wheat , Martin v. Mohrenschildt , Saeid Habibi

We analyze the Kozachenko--Leonenko (KL) nearest neighbor estimator for the differential entropy. We obtain the first uniform upper bound on its performance over H\"older balls on a torus without assuming any conditions on how close the…

Machine Learning · Statistics 2018-09-13 Jiantao Jiao , Weihao Gao , Yanjun Han

Kernel Stein discrepancies (KSDs) have emerged as a powerful tool for quantifying goodness-of-fit over the last decade, featuring numerous successful applications. To the best of our knowledge, all existing KSD estimators with known rate…

Machine Learning · Statistics 2026-03-31 Jose Cribeiro-Ramallo , Agnideep Aich , Florian Kalinke , Ashit Baran Aich , Zoltán Szabó

This paper studies the minimax rate of nonparametric conditional density estimation under a weighted absolute value loss function in a multivariate setting. We first demonstrate that conditional density estimation is impossible if one only…

Statistics Theory · Mathematics 2021-03-15 Michael Li , Matey Neykov , Sivaraman Balakrishnan

We study the convergence properties, in Hellinger and related distances, of nonparametric density estimators based on measure transport. These estimators represent the measure of interest as the pushforward of a chosen reference…

Statistics Theory · Mathematics 2022-09-20 Sven Wang , Youssef Marzouk

We consider the problem of estimating probability density functions based on sample data, using a finite mixture of densities from some component class. To this end, we introduce the $h$-lifted Kullback--Leibler (KL) divergence as a…

Machine Learning · Statistics 2024-12-24 Mark Chiu Chong , Hien Duy Nguyen , TrungTin Nguyen

The ratio of two probability densities can be used for solving various machine learning tasks such as covariate shift adaptation (importance sampling), outlier detection (likelihood-ratio test), and feature selection (mutual information).…

Machine Learning · Statistics 2009-12-16 Takafumi Kanamori , Taiji Suzuki , Masashi Sugiyama

A new maximum likelihood method for deconvoluting a continuous density with a positive lower bound on a known compact support in additive measurement error models with known error distribution using the approximate Bernstein type polynomial…

Methodology · Statistics 2018-01-30 Zhong Guan