中文

Forget the Data and Fine-Tuning! Just Fold the Network to Compress

机器学习 2025-08-13 v2 人工智能

摘要

我们提出了一种名为网络折叠(model folding)的新型数据无关模型压缩技术,通过合并跨层结构相似的神经元来显著减少模型大小,无需微调或访问训练数据。与现有方法不同,模型折叠在压缩过程中保留数据统计信息,利用 K 均值聚类,并采用新型数据无关技术防止方差塌陷或爆炸。我们的理论框架和在标准基准测试(包括 ResNet18 和 LLaMA-7B)上的实验表明,模型折叠在性能上与数据驱动压缩技术相当,且在各种稀疏度下均优于最近提出的数据无关方法,尤其是在高稀疏度水平下。这一方法在压缩大规模模型方面尤为有效,适用于资源受限环境的部署。

关键词

引用

@article{arxiv.2502.10216,
  title  = {Forget the Data and Fine-Tuning! Just Fold the Network to Compress},
  author = {Dong Wang and Haris Šikić and Lothar Thiele and Olga Saukh},
  journal= {arXiv preprint arXiv:2502.10216},
  year   = {2025}
}

备注

This paper has been accepted by The Thirteenth International Conference on Learning Representations(ICLR), 2025