Learning Distributions from Multiple Data Providers
Abstract
Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples. The goal is to learn an unknown distribution on a finite domain . The learner is given a fixed family of queryable sets , and each query to returns an independent sample from the conditional distribution . Learnability is governed by the co-occurrence graph associated with : two domain elements are adjacent if they appear together in some queryable set. Pointwise consistency is achievable when this graph is connected on the target support. PAC learning requires more: it is possible when the co-occurrence graph is complete. The optimal sample complexity of PAC learning ranges from nearly linear to quadratic. Every query family with complete co-occurrence graph admits sample complexity , and this bound is tight in the worst case. On the other hand, if is queryable then ordinary sampling improves the bound to , and this cannot be improved further even if every set is queryable. More generally, we identify hierarchical comparabilityas a sufficient structural condition on under which the optimal complexity is nearly linear, , with pairwise query families as a canonical example. Finally, the full range of polynomial rates between linear and quadratic is attainable: for every , there exists a query family with optimal PAC rate .
Cite
@article{arxiv.2607.24732,
title = {Learning Distributions from Multiple Data Providers},
author = {Jon Kleinberg and Amin Saberi and Xizhi Tan and Grigoris Velegkas},
journal= {arXiv preprint arXiv:2607.24732},
year = {2026}
}