English

DataFed: Towards Reproducible Research via Federated Data Management

Databases 2020-04-09 v1 Computers and Society

Abstract

The increasingly collaborative, globalized nature of scientific research combined with the need to share data and the explosion in data volumes present an urgent need for a scientific data management system (SDMS). An SDMS presents a logical and holistic view of data that greatly simplifies and empowers data organization, curation, searching, sharing, dissemination, etc. We present DataFed -- a lightweight, distributed SDMS that spans a federation of storage systems within a loosely-coupled network of scientific facilities. Unlike existing SDMS offerings, DataFed uses high-performance and scalable user management and data transfer technologies that simplify deployment, maintenance, and expansion of DataFed. DataFed provides web-based and command-line interfaces to manage data and integrate with complex scientific workflows. DataFed represents a step towards reproducible scientific research by enabling reliable staging of the correct data at the desired environment.

Keywords

Cite

@article{arxiv.2004.03710,
  title  = {DataFed: Towards Reproducible Research via Federated Data Management},
  author = {Dale Stansberry and Suhas Somnath and Jessica Breet and Gregory Shutt and Mallikarjun Shankar},
  journal= {arXiv preprint arXiv:2004.03710},
  year   = {2020}
}

Comments

Part of conference proceedings at the 6th Annual Conference on Computational Science & Computational Intelligence held at Las Vegas, NV, USA on Dec 05-07 2019

R2 v1 2026-06-23T14:43:35.162Z