D4M 2.0 Schema:一种用于 Accumulo 数据库的通用高性能模式
数据库
2015-05-26 v1 天体物理仪器与方法
分布式、并行与集群计算
摘要
非传统、弱一致性的三元组存储数据库是许多网络公司的支柱(例如 Google Big Table、Amazon Dynamo 和 Facebook Cassandra)。Apache Accumulo 数据库是一种高性能开源弱一致性数据库,广泛应用于政府应用。要获得 Accumulo 的全部优势,需要使用新颖的模式。动态分布式维度数据模型 (D4M)[http://d4m.mit.edu] 提供了一个基于关联数组的统一数学框架,涵盖了传统(即 SQL)和非传统数据库。对于非传统数据库,D4M 自然导出一种通用模式,可用于对数据集中的每个唯一字符串进行完全索引和快速查询。D4M 2.0 Schema 已应用于网络空间、生物信息学、科学引文、自由文本和社交媒体数据,几乎无需定制。D4M 2.0 Schema 简单,需要极少的解析,并实现了已发布的最高 Accumulo 摄入率。D4M 2.0 Schema 的优势独立于 D4M 接口;任何 Accumulo 接口均可通过使用 D4M 2.0 Schema 获得这些优势。
引用
@article{arxiv.1407.3859,
title = {D4M 2.0 Schema: A General Purpose High Performance Schema for the Accumulo Database},
author = {Jeremy Kepner and Christian Anderson and William Arcand and David Bestor and Bill Bergeron and Chansup Byun and Matthew Hubbell and Peter Michaleas and Julie Mullen and David O'Gwynn and Andrew Prout and Albert Reuther and Antonio Rosa and Charles Yee},
journal= {arXiv preprint arXiv:1407.3859},
year = {2015}
}
备注
6 pages; IEEE HPEC 2013