English
Related papers

Related papers: CMS Analysis and Data Reduction with Apache Spark

200 papers

Distributed dataflow systems like Apache Spark and Apache Hadoop enable data-parallel processing of large datasets on clusters. Yet, selecting appropriate computational resources for dataflow jobs -- that neither lead to bottlenecks nor to…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-01-11 Jonathan Will , Lauritz Thamsen , Jonathan Bader , Dominik Scheinert , Odej Kao

HEP-Frame is a new C++ package designed to efficiently perform analyses of data sets from a very large number of events, like those available at the Large Hadron Collider (LHC) at CERN, Geneva. It mainly targets high performance servers and…

High Energy Physics - Experiment · Physics 2023-03-10 A. Pereira , A. Onofre , A. Proenca

Data of the order of terabytes, petabytes, or beyond is known as Big Data. This data cannot be processed using the traditional database software, and hence there comes the need for Big Data Platforms. By combining the capabilities and…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-11-05 Tanuja Patanshetti , Ashish Anil Pawar , Disha Patel , Sanket Thakare

The massive data sets from today's particle physics experiments present a variety of challenges amenable to the tools developed by the statistics community. From the real-time decision of what subset of data to record on permanent storage,…

High Energy Physics - Experiment · Physics 2007-05-23 Bruce Knuteson , Paul Padley

Over the past decade, the fourth paradigm of data-intensive science rapidly became a major driving concept of multiple application domains encompassing and generating large-scale devices such as light sources and cutting edge telescopes.…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-06-05 Nikolay Malitsky , Ralph Castain , Matt Cowan

Big Data has become prominent throughout many scientific fields and, as a result, scientific communities have sought out Big Data frameworks to accelerate the processing of their increasingly data-intensive pipelines. However, while…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-05-31 Valerie Hayot-Sasson , Tristan Glatard

Data analysis in fundamental sciences nowadays is an essential process that pushes frontiers of our knowledge and leads to new discoveries. At the same time we can see that complexity of those analyses increases fast due to a)~enormous…

Data Analysis, Statistics and Probability · Physics 2016-01-20 Tatiana Likhomanenko , Alex Rogozhnikov , Alexander Baranov , Egor Khairullin , Andrey Ustyuzhanin

BigBench is the new standard (TPCx-BB) for benchmarking and testing Big Data systems. The TPCx-BB specification describes several business use cases -- queries -- which require a broad combination of data extraction techniques including…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-07-07 Nicolas Poggi , Alejandro Montero , David Carrera

While cluster computing frameworks are continuously evolving to provide real-time data analysis capabilities, Apache Spark has managed to be at the forefront of big data analytics for being a unified framework for both, batch and stream…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-07-31 Ahsan Javed Awan , Mats Brorsson , Vladimir Vlassov , Eduard Ayguade

The growth of big data in domains such as Earth Sciences, Social Networks, Physical Sciences, etc. has lead to an immense need for efficient and scalable linear algebra operations, e.g. Matrix inversion. Existing methods for efficient and…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-01-16 Chandan Misra , Sourangshu Bhattacharya , Soumya K. Ghosh

Real-world data from diverse domains require real-time scalable analysis. Large-scale data processing frameworks or engines such as Hadoop fall short when results are needed on-the-fly. Apache Spark's streaming library is increasingly…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-08-02 Janak Dahal , Elias Ioup , Shaikh Arifuzzaman , Mahdi Abdelguerfi

Access plan recommendation is a query optimization approach that executes new queries using prior created query execution plans (QEPs). The query optimizer divides the query space into clusters in the mentioned method. However, traditional…

Databases · Computer Science 2022-10-14 Elham Azhir , Mehdi Hosseinzadeh , Faheem Khan , Amir Mosavi

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for…

Instrumentation and Detectors · Physics 2026-03-12 Julia Gonski , Jenni Ott , Shiva Abbaszadeh , Sagar Addepalli , Matteo Cremonesi , Jennet Dickinson , Giuseppe Di Guglielmo , Erdem Yigit Ertorer , Lindsey Gray , Ryan Herbst , Christian Herwig , Tae Min Hong , Benedikt Maier , Maryam Bayat Makou , David Miller , Mark S. Neubauer , Cristián Peña , Dylan Rankin , Seon-Hee , Seo , Giordon Stark , Alexander Tapper , Audrey Corbeil Therrien , Ioannis Xiotidis , Keisuke Yoshihara , G Abarajithan , Sagar Addepalli , Nural Akchurin , Carlos Argüelles , Saptaparna Bhattacharya , Lorenzo Borella , Christian Boutan , Tom Braine , James Brau , Martin Breidenbach , Antonio Chahine , Talal Ahmed Chowdhury , Yuan-Tang Chou , Seokju Chung , Alberto Coppi , Mariarosaria D'Alfonso , Abhilasha Dave , Chance Desmet , Angela Di Fulvio , Karri DiPetrillo , Javier Duarte , Auralee Edelen , Jan Eysermans , Yongbin Feng , Emmett Forrestel , Dolores Garcia , Loredana Gastaldo , Julián García Pardiñas , Lino Gerlach , Loukas Gouskos , Katya Govorkova , Carl Grace , Christopher Grant , Philip Harris , Ciaran Hasnip , Timon Heim , Abraham Holtermann , Tae Min Hong , Gian Michele Innocenti , Koji Ishidoshiro , Miaochen Jin , Jyothisraj Johnson , Stephen Jones , Andreas Jung , Georgia Karagiorgi , Ryan Kastner , Nicholas Kamp , Doojin Kim , Kyoungchul Kong , Katie Kudela , Jelena Lalic , Bo-Cheng Lai , Yun-Tsung Lai , Tommy Lam , Jeffrey Lazar , Aobo Li , Zepeng Li , Haoyun Liu , Vladimir Lončar , Luca Macchiarulo , Christopher Madrid , Benedikt Maier , Zhenghua Ma , Prashansa Mukim , Mark S. Neubauer , Victoria Nguyen , Sungbin Oh , Isobel Ojalvo , Hideyoshi Ozaki , Simone Pagan Griso , Myeonghun Park , Christoph Paus , Santosh Parajuli , Benjamin Parpillon , Sara Pozzi , Ema Puljak , Benjamin Ramhorst , Amy Roberts , Larry Ruckman , Kate Scholberg , Sebastian Schmitt , Noah Singer , Eluned Anne Smith , Alexandre Sousa , Michael Spannowsky , Sioni Summers , Yanwen Sun , Daniel Tapia Takaki , Antonino Tumeo , Caterina Vernieri , Belina von Krosigk , Yash Vora , Linyan Wan , Michael H. L. S. Wang , Amanda Weinstein , Andy White , Simon Williams , Felix Yu

Data from high-energy physics (HEP) experiments are collected with significant financial and human effort and are mostly unique. At the same time, HEP has no coherent strategy for data preservation and re-use. An inter-experimental Study…

High Energy Physics - Experiment · Physics 2012-08-27 Dphep Study Group

In High Energy Physics (HEP), experimentalists generate large volumes of data that, when analyzed, helps us better understand the fundamental particles and their interactions. This data is often captured in many files of small size,…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-05-04 Sunwoo Lee , Kai-yuan Hou , Kewei Wang , Saba Sehrish , Marc Paterno , James Kowalkowski , Quincey Koziol , Robert Ross , Ankit Agrawal , Alok Choudhary , Wei-keng Liao

The accumulation of a large amount of new experimental data at an impressive rate at present and future collider experiments has led to important questions concerning data storage and organization, their public access and usability, as well…

High Energy Physics - Phenomenology · Physics 2019-07-30 Andrea Ceccarelli , Andrea Cioni , Maria Vittoria Garzelli , Piergiulio Lenzi , Laura Redapi

The scale of scientific High Performance Computing (HPC) and High Throughput Computing (HTC) has increased significantly in recent years, and is becoming sensitive to total energy use and cost. Energy-efficiency has thus become an important…

Distributed, Parallel, and Cluster Computing · Computer Science 2014-10-14 David Abdurachmanov , Peter Elmer , Giulio Eulisse , Robert Knight , Tapio Niemi , Jukka K. Nurminen , Filip Nyback , Goncalo Pestana , Zhonghong Ou , Kashif Khan

The IRIS-HEP Analysis Grand Challenge (AGC) is designed to be a realistic environment for investigating how analysis methods scale to the demands of the HL-LHC. The analysis task is based on publicly available Open Data and allows for…

High Energy Physics - Experiment · Physics 2023-04-12 Oksana Shadura , Alexander Held

Modern distributed data processing systems struggle to balance performance, maintainability, and developer productivity when integrating machine learning at scale. These challenges intensify in large collaborative environments due to high…

This paper presents a modern and scalable framework for analyzing Detector Control System (DCS) data from the ATLAS experiment at CERN. The DCS data, stored in an Oracle database via the WinCC OA system, is optimized for transactional…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-01-24 Luca Canali , Andrea Formica , Michelle Ann Solis