English

MixML: A Unified Analysis of Weakly Consistent Parallel Learning

Machine Learning 2020-06-09 v2 Machine Learning

Abstract

Parallelism is a ubiquitous method for accelerating machine learning algorithms. However, theoretical analysis of parallel learning is usually done in an algorithm- and protocol-specific setting, giving little insight about how changes in the structure of communication could affect convergence. In this paper we propose MixML, a general framework for analyzing convergence of weakly consistent parallel machine learning. Our framework includes: (1) a unified way of modeling the communication process among parallel workers; (2) a new parameter, the mixing time tmix, that quantifies how the communication process affects convergence; and (3) a principled way of converting a convergence proof for a sequential algorithm into one for a parallel version that depends only on tmix. We show MixML recovers and improves on known convergence bounds for asynchronous and/or decentralized versions of many algorithms, includingSGD and AMSGrad. Our experiments substantiate the theory and show the dependency of convergence on the underlying mixing time.

Keywords

Cite

@article{arxiv.2005.06706,
  title  = {MixML: A Unified Analysis of Weakly Consistent Parallel Learning},
  author = {Yucheng Lu and Jack Nash and Christopher De Sa},
  journal= {arXiv preprint arXiv:2005.06706},
  year   = {2020}
}
R2 v1 2026-06-23T15:32:05.484Z