English
Related papers

Related papers: CMS Analysis and Data Reduction with Apache Spark

200 papers

The proliferation of sensor technologies and advancements in data collection methods have enabled the accumulation of very large amounts of data. Increasingly, these datasets are considered for scientific research. However, the design of…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-07-28 Fatemeh Rouzbeh , Ananth Grama , Paul Griffin , Mohammad Adibuzzaman

Due to the significant importance of Big Data analysis, especially in business-related topics such as improving services, finding potential customers, and selecting practical approaches to manage income and expenses, many companies attempt…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-06-01 Mohammad Sina Kiarostami

Big data processing is a hot topic in today's computer science world. There is a significant demand for analysing big data to satisfy many requirements of many industries. Emergence of the Kappa architecture created a strong requirement for…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-10-17 Shelan Perera , Ashansa Perera , Kamal Hakimzadeh

With the advent of extremely high dimensional datasets, dimensionality reduction techniques are becoming mandatory. Among many techniques, feature selection has been growing in interest as an important tool to identify relevant features on…

The Large Hadron Collider (LHC) at CERN has generated in the last decade an unprecedented volume of data for the High-Energy Physics (HEP) field. Scientific collaborations interested in analysing such data very often require computing power…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-06-03 Jacek Kuśnierz , Vincenzo Eduardo Padulano , Maciej Malawski , Kamil Burkiewicz , Enric Tejedor Saavedra , Pedro Alonso-Jordá , Michael Pitt , Valentina Avati

The objective of this work was to utilize BigBench [1] as a Big Data benchmark and evaluate and compare two processing engines: MapReduce [2] and Spark [3]. MapReduce is the established engine for processing data on Hadoop. Spark is a…

Databases · Computer Science 2016-01-14 Todor Ivanov , Max-Georg Beer

As particle physics experiments push their limits on both the energy and the intensity frontiers, the amount and complexity of the produced data are also expected to increase accordingly. With such large data volumes, next-generation…

High Energy Physics - Experiment · Physics 2022-03-16 Amit Bashyal , Peter Van Gemmeren , Saba Sehrish , Kyle Knoepfel , Suren Byna , Qiao Kang

In this paper, we evaluate Apache Spark for a data-intensive machine learning problem. Our use case focuses on policy diffusion detection across the state legislatures in the United States over time. Previous work on policy diffusion has…

Computation and Language · Computer Science 2019-12-03 Alexey Svyatkovskiy , Kosuke Imai , Mary Kroeger , Yuki Shiraito

Data from particle physics experiments are unique and are often the result of a very large investment of resources. Given the potential scientific impact of these data, which goes far beyond the immediate priorities of the experimental…

High Energy Physics - Phenomenology · Physics 2025-04-02 Jon Butterworth , Sabine Kraml , Harrison Prosper , Andy Buckley , Louie Corpe , Cristinel Diaconu , Mark Goodsell , Philippe Gras , Martin Habedank , Clemens Lange , Kati Lassila-Perini , André Lessa , Rakhi Mahbubani , Judita Mamužić , Zach Marshall , Thomas McCauley , Humberto Reyes-Gonzalez , Krzysztof Rolbiecki , Sezen Sekmen , Giordon Stark , Graeme Watt , Jonas Würzinger , Shehu AbdusSalam , Aytul Adiguzel , Amine Ahriche , Ben Allanach , Mohammad M. Altakach , Jack Y. Araz , Alexandre Arbey , Saiyad Ashanujjaman , Volker Austrup , Emanuele Bagnaschi , Sumit Banik , Csaba Balazs , Daniele Barducci , Philip Bechtle , Samuel Bein , Nicolas Berger , Tisa Biswas , Fawzi Boudjema , Jamie Boyd , Carsten Burgard , Jackson Burzynski , Jordan Byers , Giacomo Cacciapaglia , Cécile Caillol , Orhan Cakir , Christopher Chang , Gang Chen , Andrea Coccaro , Yara do Amaral Coutinho , Andreas Crivellin , Leo Constantin , Giovanna Cottin , Hridoy Debnath , Mehmet Demirci , Juhi Dutta , Joe Egan , Carlos Erice Cid , Farida Fassi , Matthew Feickert , Arnaud Ferrari , Pavel Fileviez Perez , Dillon S. Fitzgerald , Roberto Franceschini , Benjamin Fuks , Lorenz Gärtner , Kirtiman Ghosh , Andrea Giammanco , Alejandro Gomez Espinosa , Letícia M. Guedes , Giovanni Guerrieri , Christian Gütschow , Abdelhamid Haddad , Mahsana Haleem , Hassane Hamdaoui , Sven Heinemeyer , Lukas Heinrich , Ben Hodkinson , Gabriela Hoff , Cyril Hugonie , Sihyun Jeon , Adil Jueid , Deepak Kar , Anna Kaczmarska , Venus Keus , Michael Klasen , Kyoungchul Kong , Joachim Kopp , Michael Krämer , Manuel Kunkel , Bertrand Laforge , Theodota Lagouri , Eric Lancon , Peilian Li , Gabriela Lima Lichtenstein , Yang Liu , Steven Lowette , Jayita Lahiri , Siddharth Prasad Maharathy , Farvah Mahmoudi , Vasiliki A. Mitsou , Sanjoy Mandal , Michelangelo Mangano , Kentarou Mawatari , Peter Meinzinger , Manimala Mitra , Mojtaba Mohammadi Najafabadi , Sahana Narasimha , Siavash Neshatpour , Jacinto P. Neto , Mark Neubauer , Mohammad Nourbakhsh , Giacomo Ortona , Rojalin Padhan , Orlando Panella , Timothée Pascal , Brian Petersen , Werner Porod , Farinaldo S. Queiroz , Shakeel Ur Rahaman , Are Raklev , Hossein Rashidi , Patricia Rebello Teles , Federico Leo Redi , Jürgen Reuter , Tania Robens , Abhishek Roy , Subham Saha , Ahmetcan Sansar , Kadir Saygin , Nikita Schmal , Jeffrey Shahinian , Sukanya Sinha , Ricardo C. Silva , Tim Smith , Tibor Šimko , Andrzej Siodmok , Ana M. Teixeira , Tamara Vázquez Schröder , Carlos Vázquez Sierra , Yoxara Villamizar , Wolfgang Waltenberger , Peng Wang , Martin White , Kimiko Yamashita , Ekin Yoruk , Xuai Zhuang

Scientific analyses commonly compose multiple single-process programs into a dataflow. An end-to-end dataflow of single-process programs is known as a many-task application. Typically, tools from the HPC software stack are used to…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-03-15 Zhao Zhang , Kyle Barbary , Frank Austin Nothaft , Evan Sparks , Oliver Zahn , Michael J. Franklin , David A. Patterson , Saul Perlmutter

We investigate the performance of Apache Spark, a cluster computing framework, for analyzing data from future LSST-like galaxy surveys. Apache Spark attempts to address big data problems have hitherto proved successful in the industry, but…

Instrumentation and Methods for Astrophysics · Physics 2018-10-17 Julien Peloton , Christian Arnault , Stéphane Plaszczynski

Each LHC experiment will produce datasets with sizes of order one petabyte per year. All of this data must be stored, processed, transferred, simulated and analyzed, which requires a computing system of a larger scale than ever mounted for…

Instrumentation and Detectors · Physics 2009-10-05 Kenneth Bloom

There is growing interest in the issues of preservation and re-use of the records of science, in the "digital era". The aim of the PARSE.Insight project, partly financed by the European Commission under the Seventh Framework Program, is…

Digital Libraries · Computer Science 2009-06-03 Andre Holzner , Peter Igo-Kemenes , Salvatore Mele

Distributed approaches based on the map-reduce programming paradigm have started to be proposed in the bioinformatics domain, due to the large amount of data produced by the next-generation sequencing techniques. However, the use of…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-07-05 Umberto Ferraro Petrillo , Mara Sorella , Giuseppe Cattaneo , Raffaele Giancarlo , Simona Rombo

This draft report summarizes and details the findings, results, and recommendations derived from the ASCR/HEP Exascale Requirements Review meeting held in June, 2015. The main conclusions are as follows. 1) Larger, more capable computing…

Every year the PHENIX collaboration deals with increasing volume of data (now about 1/4 PB/year). Apparently the more data the more questions how to process all the data in most efficient way. In recent past many developments in HEP…

Distributed, Parallel, and Cluster Computing · Computer Science 2007-05-23 Barbara Jacak , Roy Lacey , Dave Morrison , Irina Sourikova , Andrey Shevel , Qiu Zhiping

The Apache Spark framework for distributed computation is popular in the data analytics community due to its ease of use, but its MapReduce-style programming model can incur significant overheads when performing computations that do not map…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-06-06 Alex Gittens , Kai Rothauge , Shusen Wang , Michael W. Mahoney , Jey Kottalam , Lisa Gerhardt , Prabhat , Michael Ringenburg , Kristyn Maschhoff

Apache Hadoop and Spark are gaining prominence in Big Data processing and analytics. Both of them are widely deployed on Internet companies. On the other hand, high-performance data analysis requirements are causing academical and…

Performance · Computer Science 2014-03-17 Fan Liang , Chen Feng , Xiaoyi Lu , Zhiwei Xu

Data from high-energy physics (HEP) experiments are collected with significant financial and human effort and are in many cases unique. At the same time, HEP has no coherent strategy for data preservation and re-use, and many important and…

High Energy Physics - Experiment · Physics 2015-05-27 David M. South

In the big data era of observational oceanography, passive acoustics datasets are becoming too high volume to be processed on local computers due to their processor and memory limitations. As a result there is a current need for our…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-06-10 Paul Nguyen Hong Duc , Dorian Cazau