English

Meta-analysis of heterogeneous data: integrative sparse regression in high-dimensions

Methodology 2022-07-01 v2 Machine Learning

Abstract

We consider the task of meta-analysis in high-dimensional settings in which the data sources are similar but non-identical. To borrow strength across such heterogeneous datasets, we introduce a global parameter that emphasizes interpretability and statistical efficiency in the presence of heterogeneity. We also propose a one-shot estimator of the global parameter that preserves the anonymity of the data sources and converges at a rate that depends on the size of the combined dataset. For high-dimensional linear model settings, we demonstrate the superiority of our identification restrictions in adapting to a previously seen data distribution as well as predicting for a new/unseen data distribution. Finally, we demonstrate the benefits of our approach on a large-scale drug treatment dataset involving several different cancer cell-lines.

Keywords

Cite

@article{arxiv.1912.11928,
  title  = {Meta-analysis of heterogeneous data: integrative sparse regression in high-dimensions},
  author = {Subha Maity and Yuekai Sun and Moulinath Banerjee},
  journal= {arXiv preprint arXiv:1912.11928},
  year   = {2022}
}
R2 v1 2026-06-23T12:56:56.288Z