English

Parallel implementation of fast randomized algorithms for the decomposition of low rank matrices

Distributed, Parallel, and Cluster Computing 2014-04-02 v2

Abstract

We analyze the parallel performance of randomized interpolative decomposition by decomposing low rank complex-valued Gaussian random matrices up to 64 GB. We chose a Cray XMT supercomputer as it provides an almost ideal PRAM model permitting quick investigation of parallel algorithms without obfuscation from hardware idiosyncrasies. We obtain that on non-square matrices performance becomes very good, with overall runtime over 70 times faster on 128 processors. We also verify that numerically discovered error bounds still hold on matrices nearly two orders of magnitude larger than those previously tested.

Keywords

Cite

@article{arxiv.1205.3830,
  title  = {Parallel implementation of fast randomized algorithms for the decomposition of low rank matrices},
  author = {Andrew Lucas and Mark Stalzer and John Feo},
  journal= {arXiv preprint arXiv:1205.3830},
  year   = {2014}
}

Comments

9 pages, 2 figures, 5 tables. v2: extended version. this is a preprint of a published paper - see published version for definitive version

R2 v1 2026-06-21T21:05:25.135Z