English

LoRA Training in the NTK Regime has No Spurious Local Minima

Machine Learning 2024-05-29 v3 Optimization and Control

Abstract

Low-rank adaptation (LoRA) has become the standard approach for parameter-efficient fine-tuning of large language models (LLM), but our theoretical understanding of LoRA has been limited. In this work, we theoretically analyze LoRA fine-tuning in the neural tangent kernel (NTK) regime with NN data points, showing: (i) full fine-tuning (without LoRA) admits a low-rank solution of rank rNr\lesssim \sqrt{N}; (ii) using LoRA with rank rNr\gtrsim \sqrt{N} eliminates spurious local minima, allowing gradient descent to find the low-rank solutions; (iii) the low-rank solution found using LoRA generalizes well.

Keywords

Cite

@article{arxiv.2402.11867,
  title  = {LoRA Training in the NTK Regime has No Spurious Local Minima},
  author = {Uijeong Jang and Jason D. Lee and Ernest K. Ryu},
  journal= {arXiv preprint arXiv:2402.11867},
  year   = {2024}
}

Comments

23 pages