Forget the Data and Fine-Tuning! Just Fold the Network to Compress
机器学习
2025-08-13 v2 人工智能
摘要
我们提出了一种名为网络折叠(model folding)的新型数据无关模型压缩技术,通过合并跨层结构相似的神经元来显著减少模型大小,无需微调或访问训练数据。与现有方法不同,模型折叠在压缩过程中保留数据统计信息,利用 K 均值聚类,并采用新型数据无关技术防止方差塌陷或爆炸。我们的理论框架和在标准基准测试(包括 ResNet18 和 LLaMA-7B)上的实验表明,模型折叠在性能上与数据驱动压缩技术相当,且在各种稀疏度下均优于最近提出的数据无关方法,尤其是在高稀疏度水平下。这一方法在压缩大规模模型方面尤为有效,适用于资源受限环境的部署。
引用
@article{arxiv.2502.10216,
title = {Forget the Data and Fine-Tuning! Just Fold the Network to Compress},
author = {Dong Wang and Haris Šikić and Lothar Thiele and Olga Saukh},
journal= {arXiv preprint arXiv:2502.10216},
year = {2025}
}
备注
This paper has been accepted by The Thirteenth International Conference on Learning Representations(ICLR), 2025