ALinFiK:用于可扩展第三方 LLM 数据定价的线性化未来影响核学习方法
机器学习
2025-05-14 v2
摘要
大型语言模型 (LLM) 高度依赖高质量的训练数据,使得数据定价对于优化模型性能尤其在预算有限的情况下至关重要。本文旨在提供一种惠及数据提供方和模型开发者的第三方数据定价方法。我们引入了线性化未来影响核 (LinFiK),用于评估单个数据样本在训练期间提高 LLM 性能的价值。我们进一步提出了 ALinFiK,以近似 LinFiK 的学习策略,实现可扩展的数据定价。我们的全面评估表明,该方法在有效性和效率方面均优于现有基线方法,随着 LLM 参数的增加,展现出显著的可扩展性优势。
引用
@article{arxiv.2503.01052,
title = {ALinFiK: Learning to Approximate Linearized Future Influence Kernel for Scalable Third-Party LLM Data Valuation},
author = {Yanzhou Pan and Huawei Lin and Yide Ran and Jiamin Chen and Xiaodong Yu and Weijie Zhao and Denghui Zhang and Zhaozhuo Xu},
journal= {arXiv preprint arXiv:2503.01052},
year = {2025}
}
备注
Proceedings of the NAACL 2025. Keywords: Influence Function, Data Valuation, Influence Estimation. https://aclanthology.org/2025.naacl-long.589/