English

Efficient Learning of Truncated Boolean Product Distributions: Influence to the Rescue

Machine Learning 2026-07-24 v1 Data Structures and Algorithms Machine Learning

Abstract

Learning the natural parameters zRnz \in \mathbb{R}^n of discrete distributions μz\mu_z from independent samples constrained to a subset S{0,1}nS \subseteq \{0,1\}^n is a foundational challenge in high-dimensional statistics. Existing methods for efficiently estimating truncated Boolean product distributions, notably the work of [Fotakis et al' COLT'20, Algorithmica '22], require either strong local connectivity assumptions on SS -- a property denoted fatness -- or stringent anti-concentration assumptions and necessitate the total mass of the truncation set to be a constant with respect to nn. Moreover, the results in [Fotakis et al' COLT'20, Algorithmica '22] suffer from sample complexities that scale as Ω(2n)\Omega(2^n) if the mass of SS is exponentially small in nn. In this work, we circumvent these limitations by analyzing the geometry of SS under the measure μz\mu_z. We refine the existing parameter estimation guarantees under the fatness assumption, improving the prior sample complexity to O(logn/ϵ2)O( \log n / \epsilon^2) for \ell_\infty-recovery, matching the untruncated minimax rate. We further generalize fatness using the notion of influence utilized in the analysis of Boolean functions and provide sufficient conditions for efficient inference. Notably, unlike previous work, our method does not require sampling at arbitrary parameterizations of the model. Lastly, we establish a theoretical lower bound demonstrating the sample complexity exhibits an intrinsic exponential dependence on the width of the model and the minimum distance between elements in the set.

Keywords

Cite

@article{arxiv.2607.22889,
  title  = {Efficient Learning of Truncated Boolean Product Distributions: Influence to the Rescue},
  author = {Rohan Chauhan and Ioannis Panageas},
  journal= {arXiv preprint arXiv:2607.22889},
  year   = {2026}
}