English

The Architectural Implications of Facebook's DNN-based Personalized Recommendation

Distributed, Parallel, and Cluster Computing 2020-02-18 v4 Machine Learning

Abstract

The widespread application of deep learning has changed the landscape of computation in the data center. In particular, personalized recommendation for content ranking is now largely accomplished leveraging deep neural networks. However, despite the importance of these models and the amount of compute cycles they consume, relatively little research attention has been devoted to systems for recommendation. To facilitate research and to advance the understanding of these workloads, this paper presents a set of real-world, production-scale DNNs for personalized recommendation coupled with relevant performance metrics for evaluation. In addition to releasing a set of open-source workloads, we conduct in-depth analysis that underpins future system design and optimization for at-scale recommendation: Inference latency varies by 60% across three Intel server generations, batching and co-location of inferences can drastically improve latency-bounded throughput, and the diverse composition of recommendation models leads to different optimization strategies.

Keywords

Cite

@article{arxiv.1906.03109,
  title  = {The Architectural Implications of Facebook's DNN-based Personalized Recommendation},
  author = {Udit Gupta and Carole-Jean Wu and Xiaodong Wang and Maxim Naumov and Brandon Reagen and David Brooks and Bradford Cottel and Kim Hazelwood and Bill Jia and Hsien-Hsin S. Lee and Andrey Malevich and Dheevatsa Mudigere and Mikhail Smelyanskiy and Liang Xiong and Xuan Zhang},
  journal= {arXiv preprint arXiv:1906.03109},
  year   = {2020}
}

Comments

11 pages