Related papers: CMS Analysis and Data Reduction with Apache Spark

Using Big Data Technologies for HEP Analysis

The HEP community is approaching an era were the excellent performances of the particle accelerators in delivering collision at high rate will force the experiments to record a large amount of information. The growing size of the datasets…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-10-02 Matteo Cremonesi , Claudio Bellini , Bianny Bian , Luca Canali , Vasileios Dimakopoulos , Peter Elmer , Ian Fisk , Maria Girone , Oliver Gutsche , Siew-Yan Hoh , Bo Jayatilaka , Viktor Khristenko , Andrea Luiselli , Andrew Melo , Evangelos Evangelos , Dominick Olivito , Jacopo Pazzini , Jim Pivarski , Alexey Svyatkovskiy , Marco Zanetti

Big Data in HEP: A comprehensive use case study

Experimental Particle Physics has been at the forefront of analyzing the worlds largest datasets for decades. The HEP community was the first to develop suitable software and computing tools for this task. In recent times, new toolkits and…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-11-23 Oliver Gutsche , Matteo Cremonesi , Peter Elmer , Bo Jayatilaka , Jim Kowalkowski , Jim Pivarski , Saba Sehrish , Cristina Mantilla Surez , Alexey Svyatkovskiy , Nhan Tran

Exploiting Apache Spark platform for CMS computing analytics

The CERN IT provides a set of Hadoop clusters featuring more than 5 PBytes of raw storage with different open-source, user-level tools available for analytical purposes. The CMS experiment started collecting a large set of computing…

Data Analysis, Statistics and Probability · Physics 2017-11-03 Marco Meoni , Valentin Kuznetsov , Luca Menichetti , Justinas Rumševičius , Tommaso Boccali , Daniele Bonacorsi

Analyzing billion-objects catalog interactively: Apache Spark for physicists

Apache Spark is a Big Data framework for working on large distributed datasets. Although widely used in the industry, it remains rather limited in the academic community or often restricted to software engineers. The goal of this paper is…

Instrumentation and Methods for Astrophysics · Physics 2019-07-17 S. Plaszczynski , J. Peloton , C. Arnault , J. E. Campagne

Technical Report: On the Usability of Hadoop MapReduce, Apache Spark & Apache Flink for Data Science

Distributed data processing platforms for cloud computing are important tools for large-scale data analytics. Apache Hadoop MapReduce has become the de facto standard in this space, though its programming interface is relatively low-level,…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-03-30 Bilal Akil , Ying Zhou , Uwe Röhm

Machine Learning Pipelines with Modern Big Data Tools for High Energy Physics

The effective utilization at scale of complex machine learning (ML) techniques for HEP use cases poses several technological challenges, most importantly on the actual implementation of dedicated end-to-end data pipelines. A solution to…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-06-17 Matteo Migliorini , Riccardo Castellotti , Luca Canali , Marco Zanetti

Gaining insight from large data volumes with ease

Efficient handling of large data-volumes becomes a necessity in today's world. It is driven by the desire to get more insight from the data and to gain a better understanding of user trends which can be transformed into economic incentives…

Data Analysis, Statistics and Probability · Physics 2019-10-02 Valentin Kuznetsov

HEP Software Foundation Community White Paper Working Group - Data Analysis and Interpretation

At the heart of experimental high energy physics (HEP) is the development of facilities and instrumentation that provide sensitivity to new phenomena. Our understanding of nature at its most fundamental level is advanced through the…

Computational Physics · Physics 2018-04-12 Lothar Bauerdick , Riccardo Maria Bianchi , Brian Bockelman , Nuno Castro , Kyle Cranmer , Peter Elmer , Robert Gardner , Maria Girone , Oliver Gutsche , Benedikt Hegner , José M. Hernández , Bodhitha Jayatilaka , David Lange , Mark S. Neubauer , Daniel S. Katz , Lukasz Kreczko , James Letts , Shawn McKee , Christoph Paus , Kevin Pedro , Jim Pivarski , Martin Ritter , Eduardo Rodrigues , Tai Sakuma , Elizabeth Sexton-Kennedy , Michael D. Sokoloff , Carl Vuosalo , Frank Würthwein , Gordon Watts

A Big Data Analysis Framework Using Apache Spark and Deep Learning

With the spreading prevalence of Big Data, many advances have recently been made in this field. Frameworks such as Apache Hadoop and Apache Spark have gained a lot of traction over the past decades and have become massively popular,…

Databases · Computer Science 2017-11-28 Anand Gupta , Hardeo Thakur , Ritvik Shrivastava , Pulkit Kumar , Sreyashi Nag

Data Preservation in High Energy Physics

Data from high-energy physics experiments are collected with significant financial and human effort and are mostly unique. However, until recently no coherent strategy existed for data preservation and re-use, and many important and complex…

High Energy Physics - Experiment · Physics 2015-06-03 Roman Kogler , David M. South , Michael Steder

A Benchmarking Study to Evaluate Apache Spark on Large-Scale Supercomputers

As dataset sizes increase, data analysis tasks in high performance computing (HPC) are increasingly dependent on sophisticated dataflows and out-of-core methods for efficient system utilization. In addition, as HPC systems grow, memory…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-10-01 George K. Thiruvathukal , Cameron Christensen , Xiaoyong Jin , François Tessier , Venkatram Vishwanath

Toward real-time data query systems in HEP

Exploratory data analysis tools must respond quickly to a user's questions, so that the answer to one question (e.g. a visualized histogram or fit) can influence the next. In some SQL-based query systems used in industry, even very large…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-11-09 Jim Pivarski , David Lange , Thanat Jatuphattharachat

Benchmarking Apache Spark and Hadoop MapReduce on Big Data Classification

Most of the popular Big Data analytics tools evolved to adapt their working environment to extract valuable information from a vast amount of unstructured data. The ability of data mining techniques to filter this helpful information from…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-09-23 Taha Tekdogan , Ali Cakmak

Building a scalable python distribution for HEP data analysis

There are numerous approaches to building analysis applications across the high-energy physics community. Among them are Python-based, or at least Python-driven, analysis workflows. We aim to ease the adoption of a Python-based analysis…

Computational Physics · Physics 2018-04-25 David Lange

The Critical Importance of Software for HEP

Particle physics has an ambitious and broad global experimental programme for the coming decades. Large investments in building new facilities are already underway or under consideration. Scaling the present processing power and data…

High Energy Physics - Experiment · Physics 2025-06-13 HEP Software Foundation , : , Christina Agapopoulou , Claire Antel , Saptaparna Bhattacharya , Steven Gardiner , Krzysztof L. Genser , James Andrew Gooding , Alexander Held , Michel Hernandez Villanueva , Michel Jouvin , Tommaso Lari , Valeriia Lukashenko , Sudhir Malik , Alexander Moreno Briceño , Stephen Mrenna , Inês Ochoa , Joseph D. Osborn , Jim Pivarski , Alan Price , Eduardo Rodrigues , Richa Sharma , Nicholas Smith , Graeme Andrew Stewart , Anna Zaborowska , Dirk Zerwas , Maarten van Veghel

Big Data Meets HPC Log Analytics: Scalable Approach to Understanding Systems at Extreme Scale

Today's high-performance computing (HPC) systems are heavily instrumented, generating logs containing information about abnormal events, such as critical conditions, faults, errors and failures, system resource utilization, and about the…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-08-24 Byung H. Park , Saurabh Hukerikar , Ryan Adamson , Christian Engelmann

Towards Interactive, Adaptive and Result-aware Big Data Analytics

As data volumes grow across applications, analytics of large amounts of data is becoming increasingly important. Big data processing frameworks such as Apache Hadoop, Apache AsterixDB, and Apache Spark have been built to meet this demand. A…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-12-15 Avinash Kumar

A Survey on Spark Ecosystem for Big Data Processing

With the explosive increase of big data in industry and academic fields, it is necessary to apply large-scale data processing systems to analysis Big Data. Arguably, Spark is state of the art in large-scale data computing systems nowadays,…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-12-17 Shanjiang Tang , Bingsheng He , Ce Yu , Yusen Li , Kun Li

Data Preservation in High Energy Physics

Data preservation significantly increases the scientific output of high-energy physics experiments during and after data acquisition. For new and ongoing experiments, the careful consideration of long-term data preservation in the…

High Energy Physics - Experiment · Physics 2025-04-01 Alexandre Arbey , Jamie Boyd , Daniel Britzger , Concetta Cartaro , Gang Chen , Gabor David , Dmitri Denisov , Cristinel Diaconu , Dirk Duellmann , Marcus Ebert , Eckhard Elsen , Jacopo Fanini , Dillon S. Fitzgerald , Benjamin Fuks , Gerardo Ganis , Achim Geiser , Takanori Hara , Lukas Heinrich , Michael D. Hildreth , Julie M. Hogan , Henry Klest , Sabine Kraml , Eric Lançon , Clemens Lange , Kati Lassila-Perini , Sergey Levonian , Dietrich Liko , Chiara Mariotti , Zach Marshall , Thomas McCauley , François Le Diberder , Jean-Yves Le Meur , Gerald Myatt , Maxim Potekhin , Michael Roney , Pablo Saiz , Heidi Schellman , Jose Benito Gonzalez , Matthias Schröder , Ulrich Schwickerath , Tim Smith , David South , Giordon Stark , Tibor Šimko , Jan Timmermans , Andrii Verbytskyi , Arne Wiebalck , Zhiqing Zhang

Understanding the Challenges and Assisting Developers with Developing Spark Applications

To process data more efficiently, big data frameworks provide data abstractions to developers. However, due to the abstraction, there may be many challenges for developers to understand and debug the data processing code. To uncover the…

Software Engineering · Computer Science 2021-03-29 Zehao Wang