English

Enabling full-speed random access to the entire memory on the A100 GPU

Performance 2024-05-21 v1 Hardware Architecture

Abstract

We describe some features of the A100 memory architecture. In particular, we give a technique to reverse-engineer some hardware layout information. Using this information, we show how to avoid TLB issues to obtain full-speed random HBM access to the entire memory, as long as we constrain any particular thread to a reduced access window of less than 64GB.

Keywords

Cite

@article{arxiv.2405.11425,
  title  = {Enabling full-speed random access to the entire memory on the A100 GPU},
  author = {Alden Walker},
  journal= {arXiv preprint arXiv:2405.11425},
  year   = {2024}
}

Comments

6 pages, 6 figures

R2 v1 2026-06-28T16:32:08.149Z