中文

通过DValCards 实现数据估值透明化

机器学习 2025-07-31 v2

摘要

在 data-centric machine learning (ML) 兴起之后,各种 data valuation 方法被提出以 quantify each datapoint 对 desired ML model performance metrics(例如 accuracy)的贡献。除了 technical applications(例如 data cleaning、data acquisition 等)之外,有人 suggest 在 data markets context 中,data buyers 可能利用此类方法来公平 compensation data owners。我们表明 data valuation metrics 本质上存在 bias and instability under simple algorithmic design choices,这导致 technical and ethical implications。通过分析 9 个 tabular classification datasets 和 6 种 data valuation methods,我们展示了 (1) common and inexpensive data pre-processing 技术会 严重改变 estimated data values;(2) 通过 data valuation metrics 的 subsampling 可能增加 class imbalance;(3) data valuation metrics 可能 low-value underrepresented group data。因此,我们 argue in favor of increased transparency associated with data valuation in-the-wild,并 introduce novel Data Valuation Cards (DValCards) framework towards this aim。DValCards 的 proliferation 将 reduce misuse of data valuation metrics,包括 in data pricing, and build trust in responsible ML systems。

关键词

引用

@article{arxiv.2506.23349,
  title  = {A case for data valuation transparency via DValCards},
  author = {Keziah Naggita and Julienne LaChance},
  journal= {arXiv preprint arXiv:2506.23349},
  year   = {2025}
}

备注

To be published in the proceedings of the Eighth AAAI/ACM Conference on Artificial Intelligence, Ethics and Society (AIES-25)