理解多模态对比学习中数据过滤收益
机器学习
2025-12-17 v1 机器学习
摘要
现代多模态表征学习的成功依赖于互联网规模的数据集。由于大量原始网络数据的质量较低,数据 curated 已成为训练管道中的关键步骤。使用训练好的模型(即教师式过滤)进行过滤已成为一种成功解决方案,利用预训练模型计算质量分数来解释教师式过滤的经验成功。在标准双模态数据生成模型下, characterize 了过滤后对比学习的性能。设定 为 n 个配对样本中正确匹配模态的比例,我们利用线性对比学习设置显示数据过滤的可 provable 利益: 在无过滤情况下,误差被上、下界为 , 在教师式过滤情况下,误差在大 范围内上界为 ,在小 范围内上界为 。
引用
@article{arxiv.2512.14230,
title = {Understanding the Gain from Data Filtering in Multimodal Contrastive Learning},
author = {Divyansh Pareek and Sewoong Oh and Simon S. Du},
journal= {arXiv preprint arXiv:2512.14230},
year = {2025}
}
备注
40 pages, 8 figures, 1 table. This work is accepted to the Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025