高效蒸馏 LLMs 以适用于边缘应用
机器学习
2024-04-03 v1 人工智能
计算与语言
摘要
超级网络训练 LLMs 在工业应用中备受关注,因为它能够以恒定的成本产出多种不同规模/延迟的模型。我们提出了一种新方法,即称为 Multistage Low-rank Fine-tuning of Super-transformers(MLFS)用于参数高效的超级网络训练。我们展示,能够获得适用于商业边缘应用的高质量编码器模型,并且当解码器-only 模型在相当程度的压缩下仍具有抗压性时,解码器可以通过显著减少训练时间进行有效切片。
引用
@article{arxiv.2404.01353,
title = {Efficiently Distilling LLMs for Edge Applications},
author = {Achintya Kundu and Fabian Lim and Aaron Chew and Laura Wynter and Penny Chong and Rhui Dih Lee},
journal= {arXiv preprint arXiv:2404.01353},
year = {2024}
}
备注
This paper has been accepted for publication in NAACL 2024 (Industry Track)