English
Related papers

Related papers: CMS Analysis and Data Reduction with Apache Spark

200 papers

The CMS experiment at the LHC accelerator at CERN relies on its computing infrastructure to stay at the frontier of High Energy Physics, searching for new phenomena and making discoveries. Even though computing plays a significant role in…

Data Analysis, Statistics and Probability · Physics 2016-12-21 Valentin Kuznetsov , Ting Li , Luca Giommi , Daniele Bonacorsi , Tony Wildish

Optimising use of the Web (WWW) for LHC data analysis is a complex problem and illustrates the challenges arising from the integration of and computation across massive amounts of information distributed worldwide. Finding the right piece…

Instrumentation and Detectors · Physics 2014-11-18 Nigel Baker , Peter Brooks , Richard McClatchey , Zsolt Kovacs , Jean-Marie Le Goff

Computing plays an essential role in all aspects of high energy physics. As computational technology evolves rapidly in new directions, and data throughput and volume continue to follow a steep trend-line, it is important for the HEP…

Querying very large RDF data sets in an efficient manner requires a sophisticated distribution strategy. Several innovative solutions have recently been proposed for optimizing data distribution with predefined query workloads. This paper…

Databases · Computer Science 2015-07-10 Olivier Curé , Hubert Naacke , Mohamed-Amine Baazizi , Bernd Amann

Objective: To (1) demonstrate the implementation of a data science platform built on open-source technology within a large, academic healthcare system and (2) describe two computational healthcare applications built on such a platform.…

English. This document is designed to study the data structures that can be used in the Apache Spark framework and to evaluate the best performing ones to implement solutions, in particular we will evaluate advantages / disadvantages…

Databases · Computer Science 2018-10-30 Massimiliano Morrelli

Big data has found applications in multiple domains. One of the largest sources of textual big data is scientific documents and papers. Big scholarly data have been used in numerous ways to create innovative applications such as…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-11-19 Samiya Khan , Xiufeng Liu , Mansaf Alam

The IRIS-HEP software institute, as a contributor to the broader HEP Python ecosystem, is developing scalable analysis infrastructure and software tools to address the upcoming HL-LHC computing challenges with new approaches and paradigms,…

Network embedding has been widely used in social recommendation and network analysis, such as recommendation systems and anomaly detection with graphs. However, most of previous approaches cannot handle large graphs efficiently, due to that…

Social and Information Networks · Computer Science 2025-10-30 Wenqing Lin

Data from high-energy physics (HEP) experiments are collected with significant financial and human effort and are mostly unique. An inter-experimental study group on HEP data preservation and long-term analysis was convened as a panel of…

High Energy Physics - Experiment · Physics 2012-05-22 Z. Akopov , Silvia Amerio , David Asner , Eduard Avetisyan , Olof Barring , James Beacham , Matthew Bellis , Gregorio Bernardi , Siegfried Bethke , Amber Boehnlein , Travis Brooks , Thomas Browder , Rene Brun , Concetta Cartaro , Marco Cattaneo , Gang Chen , David Corney , Kyle Cranmer , Ray Culbertson , Sunje Dallmeier-Tiessen , Dmitri Denisov , Cristinel Diaconu , Vitaliy Dodonov , Tony Doyle , Gregory Dubois-Felsmann , Michael Ernst , Martin Gasthuber , Achim Geiser , Fabiola Gianotti , Paolo Giubellino , Andrey Golutvin , John Gordon , Volker Guelzow , Takanori Hara , Hisaki Hayashii , Andreas Heiss , Frederic Hemmer , Fabio Hernandez , Graham Heyes , Andre Holzner , Peter Igo-Kemenes , Toru Iijima , Joe Incandela , Roger Jones , Yves Kemp , Kerstin Kleese van Dam , Juergen Knobloch , David Kreincik , Kati Lassila-Perini , Francois Le Diberder , Sergey Levonian , Aharon Levy , Qizhong Li , Bogdan Lobodzinski , Marcello Maggi , Janusz Malka , Salvatore Mele , Richard Mount , Homer Neal , Jan Olsson , Dmitri Ozerov , Leo Piilonen , Giovanni Punzi , Kevin Regimbal , Daniel Riley , Michael Roney , Robert Roser , Thomas Ruf , Yoshihide Sakai , Takashi Sasaki , Gunar Schnell , Matthias Schroeder , Yves Schutz , Jamie Shiers , Tim Smith , Rick Snider , David M. South , Rick St. Denis , Michael Steder , Jos Van Wezel , Erich Varnes , Margaret Votava , Yifang Wang , Dennis Weygand , Vicky White , Katarzyna Wichmann , Stephen Wolbers , Masanori Yamauchi , Itay Yavin , Hans von der Schmitt

HEP data-processing software must support the disparate physics needs of many experiments. For both collider and neutrino environments, HEP experiments typically use data-processing frameworks to manage the computational complexities of…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-03-29 Christopher D. Jones , Kyle Knoepfel , Paolo Calafiura , Charles Leggett , Vakhtang Tsulaia

High Energy Physics (HEP) experiments are making increasing use of GPUs and GPU dominated High Performance Computer facilities. Both the software and hardware of these systems are rapidly evolving, creating challenges for experiments to…

High Energy Physics - Experiment · Physics 2025-05-15 Mohammad Atif , Pengfei Ding , Ka Hei Martin Kwok , Charles Leggett

The shear volumes of data generated from earth observation and remote sensing technologies continue to make major impact; leaping key geospatial applications into the dual data and compute intensive era. As a consequence, this rapid…

Computer Vision and Pattern Recognition · Computer Science 2019-08-14 Dalton Lunga , Jonathan Gerrand , Hsiuhan Lexie Yang , Christopher Layton , Robert Stewart

During the recent years, a number of efficient and scalable frequent itemset mining algorithms for big data analytics have been proposed by many researchers. Initially, MapReduce-based frequent itemset mining algorithms on Hadoop cluster…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-08-06 Pankaj Singh , Sudhakar Singh , P. K. Mishra , Rakhi Garg

The paradigm of big data is characterized by the need to collect and process data sets of great volume, arriving at the systems with great velocity, in a variety of formats. Spark is a widely used big data processing system that can be…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-12-29 Duarte M. Nascimento , Miguel Ferreira , Miguel L. Pardal

Scikit-HEP is a community-driven and community-oriented project with the goal of providing an ecosystem for particle physics data analysis in Python. Scikit-HEP is a toolset of approximately twenty packages and a few "affiliated" packages.…

The Scikit-HEP project is a community-driven and community-oriented effort with the aim of providing Particle Physics at large with a Python scientific toolset containing core and common tools. The project builds on five pillars that…

Computational Physics · Physics 2019-10-02 Eduardo Rodrigues

Supervised learning algorithms are nowadays successfully scaling up to datasets that are very large in volume, leveraging the potential of in-memory cluster-computing Big Data frameworks. Still, massive datasets with a number of…

Machine Learning · Computer Science 2018-05-11 Luca Venturini , Elena Baralis , Paolo Garza

Several important and unique experimental high-energy physics programmes at a variety of facilities are coming to an end, including those at HERA, the B-factories and the Tevatron. The wealth of physics data from these experiments is the…

High Energy Physics - Experiment · Physics 2015-06-05 David M. South

Data analytic applications built upon big data processing frameworks such as Apache Spark are an important class of applications. Many of these applications are not latency-sensitive and thus can run as batch jobs in data centers. By…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-10-03 Vicent Sanz Marco , Ben Taylor , Barry Porter , Zheng Wang
‹ Prev 1 3 4 5 6 7 10 Next ›