Efficient Learning of Truncated Boolean Product Distributions: Influence to the Rescue
Abstract
Learning the natural parameters of discrete distributions from independent samples constrained to a subset is a foundational challenge in high-dimensional statistics. Existing methods for efficiently estimating truncated Boolean product distributions, notably the work of [Fotakis et al' COLT'20, Algorithmica '22], require either strong local connectivity assumptions on -- a property denoted fatness -- or stringent anti-concentration assumptions and necessitate the total mass of the truncation set to be a constant with respect to . Moreover, the results in [Fotakis et al' COLT'20, Algorithmica '22] suffer from sample complexities that scale as if the mass of is exponentially small in . In this work, we circumvent these limitations by analyzing the geometry of under the measure . We refine the existing parameter estimation guarantees under the fatness assumption, improving the prior sample complexity to for -recovery, matching the untruncated minimax rate. We further generalize fatness using the notion of influence utilized in the analysis of Boolean functions and provide sufficient conditions for efficient inference. Notably, unlike previous work, our method does not require sampling at arbitrary parameterizations of the model. Lastly, we establish a theoretical lower bound demonstrating the sample complexity exhibits an intrinsic exponential dependence on the width of the model and the minimum distance between elements in the set.
Keywords
Cite
@article{arxiv.2607.22889,
title = {Efficient Learning of Truncated Boolean Product Distributions: Influence to the Rescue},
author = {Rohan Chauhan and Ioannis Panageas},
journal= {arXiv preprint arXiv:2607.22889},
year = {2026}
}