中文
相关论文

相关论文: Integrating R and Hadoop for Big Data Analysis

200 篇论文

MapReduce, the popular programming paradigm for large-scale data processing, has traditionally been deployed over tightly-coupled clusters where the data is already locally available. The assumption that the data and compute resources are…

分布式、并行与集群计算 · 计算机科学 2012-07-31 Benjamin Heintz , Abhishek Chandra , Ramesh K. Sitaraman

Workload consolidation, sharing physical resources among multiple workloads, is a promising technique to save cost and energy in cluster computing systems. This paper highlights a few challenges of workload consolidation for Hadoop as one…

分布式、并行与集群计算 · 计算机科学 2016-11-15 Reza Moraveji , Javid Taheri , MohammadReza HosseinyFarahabady , Nikzad Babaii Rizvandi , Albert Y. Zomaya

When processing large medical imaging studies, adopting high performance grid computing resources rapidly becomes important. We recently presented a "medical image processing-as-a-service" grid framework that offers promise in utilizing the…

分布式、并行与集群计算 · 计算机科学 2017-12-27 Shunxing Bao , Yuankai Huo , Prasanna Parvathaneni , Andrew J. Plassard , Camilo Bermudez , Yuang Yao , Ilwoo Llyu , Aniruddha Gokhale , Bennett A. Landman

While advanced analysis of large dataset is in high demand, data sizes have surpassed capabilities of conventional software and hardware. Hadoop framework distributes large datasets over multiple commodity servers and performs parallel…

分布式、并行与集群计算 · 计算机科学 2015-11-17 Woo-Hyun Lee , Hee-Gook Jun , Hyoung-Joo Kim

Applying popular machine learning algorithms to large amounts of data raised new challenges for the ML practitioners. Traditional ML libraries does not support well processing of huge datasets, so that new approaches were needed.…

分布式、并行与集群计算 · 计算机科学 2016-03-30 Daniel Pop

An increasing volume of studies utilize geocomputation methods in large spatial data. There is a bottleneck in scalable computation for general scientific use as the existing solutions require high-performance computing domain knowledge and…

分布式、并行与集群计算 · 计算机科学 2025-03-06 Insang Song , Kyle P. Messier

Document clustering is a traditional, efficient and yet quite effective, text mining technique when we need to get a better insight of the documents of a collection that could be grouped together. The K-Means algorithm and the Hierarchical…

分布式、并行与集群计算 · 计算机科学 2021-12-02 Sergios Gerakidis , Sofia Megarchioti , Basilis Mamalis

With the spreading prevalence of Big Data, many advances have recently been made in this field. Frameworks such as Apache Hadoop and Apache Spark have gained a lot of traction over the past decades and have become massively popular,…

数据库 · 计算机科学 2017-11-28 Anand Gupta , Hardeo Thakur , Ritvik Shrivastava , Pulkit Kumar , Sreyashi Nag

The growing complexity and variety of Big Data platforms makes it both difficult and time consuming for all system users to properly setup and operate the systems. Another challenge is to compare the platforms in order to choose the most…

分布式、并行与集群计算 · 计算机科学 2015-10-28 Todor Ivanov , Sead Izberovic

Storage systems are essential building blocks for cloud computing infrastructures. Although high performance storage servers are the ultimate solution for cloud storage, the implementation of inexpensive storage system remains an open…

分布式、并行与集群计算 · 计算机科学 2011-12-30 Julia Myint , Thinn Thu Naing

Many relevant applications in the environmental and socioeconomic sciences use areal data, such as biodiversity checklists, agricultural statistics, or socioeconomic surveys. For applications that surpass the spatial, temporal or thematic…

数据库 · 计算机科学 2020-07-15 Steffen Ehrmann , Ralf Seppelt , Carsten Meyer

Nowadays many companies have available large amounts of raw, unstructured data. Among Big Data enabling technologies, a central place is held by the MapReduce framework and, in particular, by its open source implementation, Apache Hadoop.…

分布式、并行与集群计算 · 计算机科学 2017-01-18 Eugenio Gianniti , Danilo Ardagna , Michele Ciavotta , Mauro Passacantando

With the overwhelming amount of complex and heterogeneous data pouring from any-where, any-time, and any-device, there is undeniably an era of Big Data. The emergence of the Big Data as a disruptive technology for next generation of…

数据库 · 计算机科学 2019-03-01 Ravi Ranjan , Aditi Sharma

Apache Hadoop and Spark are gaining prominence in Big Data processing and analytics. Both of them are widely deployed on Internet companies. On the other hand, high-performance data analysis requirements are causing academical and…

性能 · 计算机科学 2014-03-17 Fan Liang , Chen Feng , Xiaoyi Lu , Zhiwei Xu

This tutorial presents a recipe for the construction of a compute cluster for processing large volumes of data, using cheap, easily available personal computer hardware (Intel/AMD based PCs) and freely available open source software (Ubuntu…

分布式、并行与集群计算 · 计算机科学 2009-12-01 Jochen L. Leidner , Gary Berosik

This chapter introduces the state-of-the-art in the emerging area of combining High Performance Computing (HPC) with Big Data Analysis. To understand the new area, the chapter first surveys the existing approaches to integrating HPC with…

分布式、并行与集群计算 · 计算机科学 2021-01-01 Yuankun Fu , Fengguang Song

This paper describes an automated approach to handling Big Data workloads on HPC systems. We describe a solution that dynamically creates a unified cluster based on YARN in an HPC Environment, without the need to configure and allocate a…

分布式、并行与集群计算 · 计算机科学 2015-07-01 Sidharth N. Kashyap , Ade J. Fewings , Jay Davies , Ian Morris , Andrew Thomas Thomas Green , Martyn F. Guest

Modern applications can generate a large amount of data from different sources with high velocity, a combination that is difficult to store and process via traditional tools. Hadoop is one framework that is used for the parallel processing…

分布式、并行与集群计算 · 计算机科学 2023-09-29 Rana Ghazali , Sahar Adabi , Ali Rezaee , Douglas G. Down , Ali Movaghar

This article dwells on the basic characteristic features of the Big Data technologies. It is analyzed the existing definition of the "big data" term. The article proposes and describes the elements of the generalized formal model of big…

数据库 · 计算机科学 2019-05-09 Shakhovska Nataliya , Veres Oleh , Hirnyak Mariia

We propose an architecture for analysing database connection logs across different instances of databases within an intranet comprising over 10,000 users and associated devices. Our system uses Flume agents to send notifications to a Hadoop…

分布式、并行与集群计算 · 计算机科学 2018-12-04 Swapneel Mehta , Prasanth Kothuri , Daniel Lanza Garcia