数据基元:全面理解大数据与 AI 工作负载的视角
分布式、并行与集群计算
2018-08-28 v1 性能
摘要
大数据与 AI 工作负载的复杂性和多样性使得理解它们变得困难且具有挑战性。本文提出了一种建模和表征大数据与 AI 工作负载的新方法。我们将每个大数据与 AI 工作负载视为在不同初始或中间数据输入上执行的一类或多类计算单元组成的流水线。每一类计算单元捕捉了共性需求,同时又合理地脱离了具体实现,因此我们称之为数据基元(data motif)。我们首次在各种各样的大数据与 AI 工作负载中识别出占据这些工作负载绝大部分运行时间的八个数据基元,包括 Matrix、Sampling、Logic、Transform、Set、Graph、Sort 和 Statistic。我们在不同软件栈上实现了这八个数据基元,作为开源大数据与 AI 基准测试套件 BigDataBench 4.0(可从 http://prof.ict.ac.cn/BigDataBench 公开获取)的微基准,并从数据大小、类型、来源和模式的角度对这些数据基元进行了全面表征,以此作为全面理解大数据与 AI 工作负载的视角。我们相信这八个数据基元不仅对于大数据与 AI 基准测试,而且对于领域特定的硬件和软件协同设计,都是很有前景的抽象和工具。
引用
@article{arxiv.1808.08512,
title = {Data Motifs: A Lens Towards Fully Understanding Big Data and AI Workloads},
author = {Wanling Gao and Jianfeng Zhan and Lei Wang and Chunjie Luo and Daoyi Zheng and Fei Tang and Biwei Xie and Chen Zheng and Xu Wen and Xiwen He and Hainan Ye and Rui Ren},
journal= {arXiv preprint arXiv:1808.08512},
year = {2018}
}
备注
The paper will be published on The 27th International Conference on Parallel Architectures and Compilation Techniques (PACT18)