天文学前身数据序列化方法的评估
摘要
来自天文管线的前身数据在建立数据处理和产品的可信度与可重复性方面起着关键作用。此外,天文学家可以查询其前身数据以回答涉及异常检测、推荐和预测等领域的问题。下一代如 Vera Rubin 观测站或 Square Kilometre Array 这样的天文巡天望远镜,能够产生 Peta 级至 Exabyte 级数据,从而放大了对前身存储或查询效率甚至微小改进的重要性。为了确定天文学家应如何存储和查询其前身数据,本文报告了turtle 和 JSON 前身序列化方法之间的比较。选取代表性数据库管理系统(DBMS)Apache Jena Fuseki 和图数据库系统 Neo4j 分别用于turtle 和 JSON。将模拟前身数据上传至和查询每个 DBMS,并测量查询准确性和耗时以及数据上传耗时作为比较指标。发现两种序列化方法都适用于此目的,两者在查询准确性方面相似。turtle 前身在存储和上传数据方面更高效。 Regarding queries, for small datasets (5MB) and simple information retrieval queries, the turtle serialisation was also found to be more efficient. However, queries for JSON serialised provenance were found to be more efficient for more complex queries which involved matching patterns across the DBMS, this effect scaled with the size of the queried provenance.
引用
@article{arxiv.2407.14290,
title = {Evaluation of Provenance Serialisations for Astronomical Provenance},
author = {Michael A. C. Johnson and Marcus Paradies and Hans-Rainer Klöckner and Albina Muzafarova and Kristen Lackeos and David J. Champion and Marta Dembska and Sirko Schindler},
journal= {arXiv preprint arXiv:2407.14290},
year = {2024}
}
备注
9 pages, 8 figures, to be published in the 16th International Workshop on Theory and Practice of Provenance