中文
相关论文

相关论文: Integrating R and Hadoop for Big Data Analysis

200 篇论文

Data integration is one of the main problems in distributed data sources. An approach is to provide an integrated mediated schema for various data sources. This research work aims at developing a framework for defining an integrated schema…

数据库 · 计算机科学 2012-11-28 Amineh Amini , Hadi Saboohi , Nasser Nemat bakhsh

Spark is an in-memory analytics platform that targets commodity server environments today. It relies on the Hadoop Distributed File System (HDFS) to persist intermediate checkpoint states and final processing results. In Spark, immutable…

分布式、并行与集群计算 · 计算机科学 2017-08-22 Mijung Kim , Jun Li , Haris Volos , Manish Marwah , Alexander Ulanov , Kimberly Keeton , Joseph Tucek , Lucy Cherkasova , Le Xu , Pradeep Fernando

In High Energy Physics (HEP), experimentalists generate large volumes of data that, when analyzed, helps us better understand the fundamental particles and their interactions. This data is often captured in many files of small size,…

分布式、并行与集群计算 · 计算机科学 2022-05-04 Sunwoo Lee , Kai-yuan Hou , Kewei Wang , Saba Sehrish , Marc Paterno , James Kowalkowski , Quincey Koziol , Robert Ross , Ankit Agrawal , Alok Choudhary , Wei-keng Liao

In the age of big data, it is important for primary research data to follow the FAIR principles of findability, accessibility, interoperability, and reusability. Data harmonization enhances interoperability and reusability by aligning…

数据库 · 计算机科学 2025-03-26 Jimmy K. Yu , Marcos Martínez-Romero , Matthew Horridge , Mete U. Akdogan , Mark A. Musen

During the process of citation matching links from bibliography entries to referenced publications are created. Such links are indicators of topical similarity between linked texts, are used in assessing the impact of the referenced…

信息检索 · 计算机科学 2013-03-28 Mateusz Fedoryszak , Dominika Tkaczyk , Łukasz Bolikowski

Complex networks are relational data sets commonly represented as graphs. The analysis of their intricate structure is relevant to many areas of science and commerce, and data sets may reach sizes that require distributed storage and…

分布式、并行与集群计算 · 计算机科学 2016-01-05 Jannis Koch , Christian L. Staudt , Maximilian Vogel , Henning Meyerhenke

This paper presents a novel high speed clustering scheme for high dimensional data streams. Data stream clustering has gained importance in different applications, for example, in network monitoring, intrusion detection, and real-time…

数据库 · 计算机科学 2015-10-13 Irshad Ahmed , Irfan Ahmed , Waseem Shahzad

To accommodate the needs of large-scale distributed P2P systems, scalable data management strategies are required, allowing applications to efficiently cope with continuously growing, highly dis tributed data. This paper addresses the…

分布式、并行与集群计算 · 计算机科学 2009-09-30 Bogdan Nicolae , Gabriel Antoniu , Luc Bougé

Cloud computing has demonstrated that processing very large datasets over commodity clusters can be done simply given the right programming model and infrastructure. In this paper, we describe the design and implementation of the Sector…

分布式、并行与集群计算 · 计算机科学 2009-01-17 Yunhong Gu , Robert L Grossman

Big Data, Cloud computing, Cloud Database Management techniques, Data Science and many more are the fantasizing words which are the future of IT industry. For all the new techniques one common thing is that they deal with Data, not just…

分布式、并行与集群计算 · 计算机科学 2016-03-29 Shweta Malhotra , Mohammad Najmud Doja , Bashir Alam , Mansaf Alam

R is a robust open-source programming language mainly used for statistical computing . Many areas of statistical research are experiencing rapid growth in the size of data sets. Methodological advances drive increased use of simulations. A…

编程语言 · 计算机科学 2019-04-10 Rahim K. Charania

The programming paradigm Map-Reduce and its main open-source implementation, Hadoop, have had an enormous impact on large scale data processing. Our goal in this expository writeup is two-fold: first, we want to present some complexity…

分布式、并行与集群计算 · 计算机科学 2012-11-29 Ashish Goel , Kamesh Munagala

The objective of our paper is to propose a Cloud computing framework which is feasible and necessary for handling huge data. In our prototype system we considered national ID database structure of Bangladesh which is prepared by election…

分布式、并行与集群计算 · 计算机科学 2014-05-21 Narzu Tarannum , Nova Ahmed

Distributed Data Processing Platforms (e.g., Hadoop, Spark, and Flink) are widely used to store and process data in a cloud environment. These platforms distribute the storage and processing of data among the computing nodes of a cloud. The…

分布式、并行与集群计算 · 计算机科学 2023-12-08 Isuru Dharmadasa , Faheem Ullah

Born in the late 20s, R is one of the most popular software for statistical computing and graphics. With the development of information technology and the advent of the big data era, great changes have taken place in the R ecosystem. Based…

其他统计学 · 统计学 2026-05-19 Tian-Yuan Huang , Zhilan Lou

The synthpop package for R https://www.synthpop.org.uk provides tools to allow data custodians to create synthetic versions of confidential microdata that can be distributed with fewer restrictions than the original. The synthesis can be…

统计计算 · 统计学 2021-11-16 Gillian M Raab , Beata Nowok , Chris Dibben

Hyperspectral remote sensing is a promising tool for a variety of applications including ecology, geology, analytical chemistry and medical research. This article presents the new \hsdar package for R statistical software, which performs a…

Data-based classification is fundamental to most branches of science. While recent years have brought enormous progress in various areas of statistical computing and clustering, some general challenges in clustering remain: model selection,…

人工智能 · 计算机科学 2007-06-13 Jens Oehlschlägel

In Smart Grid applications, as the number of deployed electric smart meters increases, massive amounts of valuable meter data is generated and collected every day. To enable reliable data collection and make business decisions fast, high…

数据库 · 计算机科学 2014-07-10 Yue Liu , Songlin Hu , Tilmann Rabl , Wantao Liu , Hans-Arno Jacobsen , Kaifeng Wu , Jian Chen , Jintao Li

Many organizations routinely analyze large datasets using systems for distributed data-parallel processing and clusters of commodity resources. Yet, users need to configure adequate resources for their data processing jobs. This requires…

分布式、并行与集群计算 · 计算机科学 2022-06-02 Lauritz Thamsen , Dominik Scheinert , Jonathan Will , Jonathan Bader , Odej Kao