极端规模湍流数据集的智能抽样:用于准确高效时空模型训练
机器学习
2025-10-27 v3 人工智能
分布式、并行与集群计算
摘要
随着摩尔定律和戴纳定律的结束,高效训练越来越需要重新思考数据volume。我们能否通过智能抽样以更少的数据训练更好的模型?为探索这一问题,我们开发了SICKLE—a一个稀疏智能策展框架,用于高效学习,特点是新颖的最大熵(MaxEnt)抽样方法、可扩展训练和能源基准测试。我们将MaxEnt与随机抽样和相空间抽样在大型直接数值模拟(DNS)湍流数据集上进行比较。在Frontier上进行大规模评估SICKLE后,我们证明,将抽样作为预处理步骤在许多情况下可以提高模型准确性并大幅降低能源消耗,观察到最高可达38倍的能源消耗降低。
引用
@article{arxiv.2508.03872,
title = {Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training},
author = {Wesley Brewer and Murali Meena Gopalakrishnan and Matthias Maiterth and Aditya Kashi and Jong Youl Choi and Pei Zhang and Stephen Nichols and Riccardo Balin and Miles Couchman and Stephen de Bruyn Kops and P. K. Yeung and Daniel Dotson and Rohini Uma-Vaideswaran and Sarp Oral and Feiyi Wang},
journal= {arXiv preprint arXiv:2508.03872},
year = {2025}
}
备注
13 pages, 9 figures, 2 tables