English

Integrating R and Hadoop for Big Data Analysis

Distributed, Parallel, and Cluster Computing 2018-02-09 v1

Abstract

Analyzing and working with big data could be very diffi cult using classical means like relational database management systems or desktop software packages for statistics and visualization. Instead, big data requires large clusters with hundreds or even thousands of computing nodes. Offi cial statistics is increasingly considering big data for deriving new statistics because big data sources could produce more relevant and timely statistics than traditional sources. One of the software tools successfully and wide spread used for storage and processing of big data sets on clusters of commodity hardware is Hadoop. Hadoop framework contains libraries, a distributed fi le-system (HDFS), a resource-management platform and implements a version of the MapReduce programming model for large scale data processing. In this paper we investigate the possibilities of integrating Hadoop with R which is a popular software used for statistical computing and data visualization. We present three ways of integrating them: R with Streaming, Rhipe and RHadoop and we emphasize the advantages and disadvantages of each solution.

Keywords

Cite

@article{arxiv.1407.4908,
  title  = {Integrating R and Hadoop for Big Data Analysis},
  author = {Bogdan Oancea and Raluca Mariana Dragoescu},
  journal= {arXiv preprint arXiv:1407.4908},
  year   = {2018}
}

Comments

Romanian Statistical Review no. 2 / 2014

R2 v1 2026-06-22T05:07:16.317Z