MetaHQ:公共组学样本与研究的和谐化高质量元数据注释
基因组学
2026-03-16 v2
摘要
基因表达库(Gene Expression Omnibus)和序列读档库(Sequence Read Archive)等公共组学数据库提供了解决新型医学问题的数据重用的广阔机会。然而,由于描述为自由文本的元数据且缺乏标准化注释,仍难以找到感兴趣的样本和研究。为此,多个研究小组对大量数据进行整理工作,以添加标准化注释,但这些注释跨越在线资源,存储格式受不同标准化准则影响,阻碍了跨来源的注释集成。我们开发了 MetaHQ 来整合和分发公共组学样本的标准化元数据。MetaHQ 包括一个包含来自 13 个来源、近 20 万 条注释的数据库,以及用于查询数据库并检索注释的友好命令行界面(CLI)。MetaHQ CLI 作为 Python 包部署在 PyPI 上(https://pypi.org/project/metahq-cli),访问 MetaHQ 数据库(https://doi.org/10.5281/zenodo.18462463)。项目源代码与文档可在 https://github.com/krishnanlab/meta-hq 获得。
引用
@article{arxiv.2602.07805,
title = {MetaHQ: Harmonized, high-quality metadata annotations of public omics samples and studies},
author = {Parker Hicks and Lydia E Valtadoros and Christopher A Mancuso and Faisal Alquadoomi and Kayla A Johnson and Sneha Sundar and Arjun Krishnan},
journal= {arXiv preprint arXiv:2602.07805},
year = {2026}
}
备注
7 pages main text, 4 pages Supplemental Figures, 1 page Supplemental Table, 1 page Supplemental File. The replacement added three references that were missing in the original submission and made minor formatting changes