面向大型科学数据集的自监督相似度搜索
天体物理仪器与方法
2021-12-02 v2 星系天体物理
计算机视觉与模式识别
摘要
我们展示了利用自监督学习来探索和利用大型无标签数据集。聚焦于来自暗能量光谱仪(DESI)遗产成像巡天最新数据发布的 4200 万张星系图像,我们首先训练一个自监督模型,以提取对每张图像中的对称性、不确定性和噪声具有鲁棒性的低维表示。然后我们利用这些表示构建并公开发布一个交互式语义相似度搜索工具。我们展示了我们的工具如何可用于在仅给定一个样本的情况下快速发现稀有天体、提高众包活动的速度,以及构建和改进用于监督式应用的训练集。虽然我们聚焦于巡天图像,但该技术可直接应用于任意维度的任何科学数据集。相似度搜索 Web 应用可在 https://github.com/georgestein/galaxy_search 找到。
引用
@article{arxiv.2110.13151,
title = {Self-supervised similarity search for large scientific datasets},
author = {George Stein and Peter Harrington and Jacqueline Blaum and Tomislav Medan and Zarija Lukic},
journal= {arXiv preprint arXiv:2110.13151},
year = {2021}
}
备注
5 pages, 2 figures. The similarity search web app can be found at https://github.com/georgestein/galaxy_search. Accepted to the Fourth Workshop on Machine Learning and the Physical Sciences (NeurIPS 2021). ArXiv admin note: text overlap with arXiv:2110.00023