Unique scientific instruments designed and operated by large global collaborations are expected to produce Exabyte-scale data volumes per year by 2030. These collaborations depend on globally distributed storage and compute to turn raw data into science. While all of these infrastructures have batch scheduling capabilities to share compute, Research and Education networks lack those capabilities. There is thus uncontrolled competition for bandwidth between and within collaborations. As a result, data "hogs" disk space at processing facilities for much longer than it takes to process, leading to vastly over-provisioned storage infrastructures. Integrated co-scheduling of networks as part of high-level managed workflows might reduce these storage needs by more than an order of magnitude. This paper describes such a solution, demonstrates its functionality in the context of the Large Hadron Collider (LHC) at CERN, and presents the next-steps towards its use in production.
@article{arxiv.2209.13714,
title = {Managed Network Services for Exascale Data Movement Across Large Global Scientific Collaborations},
author = {Frank Würthwein and Jonathan Guiang and Aashay Arora and Diego Davila and John Graham and Dima Mishin and Thomas Hutton and Igor Sfiligoi and Harvey Newman and Justas Balcas and Tom Lehman and Xi Yang and Chin Guok},
journal= {arXiv preprint arXiv:2209.13714},
year = {2023}
}
Comments
Submitted to the proceedings of the XLOOP workshop held in conjunction with Supercomputing 22