Active Linear Regression for $\ell_p$ Norms and Beyond
Abstract
We study active sampling algorithms for linear regression, which aim to query only a few entries of a target vector and output a near minimizer to , for a design matrix and loss . For norm regression for any , we give an algorithm based on Lewis weight sampling outputting a -approximate solution using just queries to for , queries for , and queries for . For , our bounds are optimal up to log factors, settling the query complexity for this range. For , our dependence on is optimal, while our dependence on is off by at most , up to log factors. Our result resolves an open question of [CD21], who gave near optimal bounds for the norm, but required samples for regression with , and gave no bounds for or . We also give the first total sensitivity bound of for loss functions of degree polynomial growth, improving a result of [TMF20]. By combining this with our techniques for regression, we obtain an active regression algorithm making queries for such loss functions, including the Tukey and Huber losses, answering another question of [CD21]. For the Huber loss, we further improve our bound to samples. Our sensitivity bounds also have many applications, including Orlicz norm subspace embeddings, robust subspace approximation, and dimension reduction for smoothed -norms. Finally, our active sampling results give the first sublinear time algorithms for Kronecker product regression under every norm.
Cite
@article{arxiv.2111.04888,
title = {Active Linear Regression for $\ell_p$ Norms and Beyond},
author = {Cameron Musco and Christopher Musco and David P. Woodruff and Taisuke Yasuda},
journal= {arXiv preprint arXiv:2111.04888},
year = {2022}
}
Comments
Abstract shortened to meet arXiv limits; v2: improved bounds; v3: improved bounds; v4: to appear in FOCS 2022