Smoothed Analysis of Learning from Positive Samples
Abstract
Binary classification from positive-only samples is a variant of PAC learning where the learner receives i.i.d. positive samples and aims to learn a classifier with low error. Previous work by Natarajan, Gereb-Graus, and Shvaytser characterized learnability and revealed a largely negative picture: almost no interesting classes, including two-dimensional halfspaces, are learnable. This poses a challenge for applications from bioinformatics to ecology, where practitioners rely on heuristics. In this work, we initiate a smoothed analysis of positive-only learning. We assume samples from a reference distribution such that the true distribution is smooth with respect to it. In stark contrast to the worst-case setting, we show that all VC classes become learnable in the smoothed model, requiring positive samples for classification error. We also give an efficient algorithm for any class admitting -approximation by degree- polynomials whose range is lower-bounded by a constant with respect to in L1-norm. It runs in time , qualitatively matching L1-regression. Our results also imply faster or more general algorithms for: (1) estimation with unknown-truncation, giving the first polynomial-time algorithm for estimating exponential-family parameters from samples truncated to an unknown set approximable by non-negative polynomials in L1 norm, improving on [KTZ FOCS19; LMZ FOCS24], who required strong L2-approximation; (2) truncation detection for broad classes, including non-product distributions, improving on [DLNS STOC24]'s who required product distributions; and (3) learning from a list of reference distributions, where samples come from distributions, one of which witnesses smoothness of , as arises when list-decoding algorithms learn samplers for from corrupted data.
Cite
@article{arxiv.2504.10428,
title = {Smoothed Analysis of Learning from Positive Samples},
author = {Jane H. Lee and Anay Mehrotra and Manolis Zampetakis},
journal= {arXiv preprint arXiv:2504.10428},
year = {2026}
}
Comments
Accepted for presentation at the 58th ACM Symposium on Theory of Computing (STOC), 2026; abstract shortened for arXiv