English

High performance on-demand de-identification of a petabyte-scale medical imaging data lake

Distributed, Parallel, and Cluster Computing 2020-08-06 v1 Performance

Abstract

With the increase in Artificial Intelligence driven approaches, researchers are requesting unprecedented volumes of medical imaging data which far exceed the capacity of traditional on-premise client-server approaches for making the data research analysis-ready. We are making available a flexible solution for on-demand de-identification that combines the use of mature software technologies with modern cloud-based distributed computing techniques to enable faster turnaround in medical imaging research. The solution is part of a broader platform that supports a secure high performance clinical data science platform.

Keywords

Cite

@article{arxiv.2008.01827,
  title  = {High performance on-demand de-identification of a petabyte-scale medical imaging data lake},
  author = {Joseph Mesterhazy and Garrick Olson and Somalee Datta},
  journal= {arXiv preprint arXiv:2008.01827},
  year   = {2020}
}

Comments

11 pages, 3 figures, 2 tables