大语言模型中的语义保留与极端压缩:我们能同时实现两者吗?
计算与语言
2025-05-13 v1 人工智能
机器学习
摘要
大语言模型 (LLM) 部署的指数增长加剧了高效模型压缩技术以减少计算和内存成本的需求。尽管剪枝和量化显示出前景,但它们的组合潜力仍 largely 未被探索。在本文中,我们检验了联合压缩及其 strategically 组合剪枝和量化是否能产生优于单一方法的性能-压缩比率。鉴于准确评估 LLM 性能的挑战,我们解决了前一评估框架的关键局限性,提出了语义保留压缩率 (SrCr) 这一新指标,用于量化模型压缩与语义保持之间的 trade-off,促进剪枝-量化配置的优化。实验表明,我们推荐的组合方式在相同理论压缩率下,平均比仅量化模型提高 20% 的性能。
引用
@article{arxiv.2505.07289,
title = {Semantic Retention and Extreme Compression in LLMs: Can We Have Both?},
author = {Stanislas Laborde and Martin Cousseau and Antoun Yaacoub and Lionel Prevost},
journal= {arXiv preprint arXiv:2505.07289},
year = {2025}
}
备注
Accepted for publication in the Proceedings of the 2025 International Joint Conference on Neural Networks (IJCNN); this arXiv version includes an appendix with 6 result tables; 10 pages, 15 figures, 7 tables