云上大数据存储系统的高效支持
分布式、并行与集群计算
2014-12-01 v1
摘要
由于相较于传统数据中心的优势,云基础设施的使用迅速增长。这些包括公有云(例如 Amazon EC2)或使用 OpenStack 部署的私有云。许多知名基础设施(例如 OpenStack 和 CloudStack)的一个共同因素是使用网络存储来存储持久化数据。然而,传统的大数据系统(包括 Hadoop)出于高性能和低成本的原因将数据存储在商用本地存储中。我们提出了一种利用本地存储在 OpenStack 上支持 Hadoop 的架构。随后,我们通过在 OpenStack 和 Amazon 上的基准测试表明,对于支持 Hadoop,本地存储具有更好的性能和更低的成本。我们得出结论,云系统应支持将本地存储用于持久化数据(除了网络存储之外),以便为 Hadoop 和其他大数据系统提供高效支持。
引用
@article{arxiv.1411.7507,
title = {Efficient Support of Big Data Storage Systems on the Cloud},
author = {Akshay MS and Suhas Mohan and Vincent Kuri and Dinkar Sitaram and H. L. Phalachandra},
journal= {arXiv preprint arXiv:1411.7507},
year = {2014}
}
备注
Presented at 2nd International Workshop on Cloud Computing Applications (ICWA) during IEEE International Conference on High Performance Computing (HiPC) 2013