English

QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning

Machine Learning 2025-08-11 v1 Computation and Language

Abstract

Finetuning large language models requires huge GPU memory, restricting the choice to acquire Larger models. While the quantized version of the Low-Rank Adaptation technique, named QLoRA, significantly alleviates this issue, finding the efficient LoRA rank is still challenging. Moreover, QLoRA is trained on a pre-defined rank and, therefore, cannot be reconfigured for its lower ranks without requiring further fine-tuning steps. This paper proposes QDyLoRA -Quantized Dynamic Low-Rank Adaptation-, as an efficient quantization approach for dynamic low-rank adaptation. Motivated by Dynamic LoRA, QDyLoRA is able to efficiently finetune LLMs on a set of pre-defined LoRA ranks. QDyLoRA enables fine-tuning Falcon-40b for ranks 1 to 64 on a single 32 GB V100-GPU through one round of fine-tuning. Experimental results show that QDyLoRA is competitive to QLoRA and outperforms when employing its optimal rank.

Keywords

Cite

@article{arxiv.2402.10462,
  title  = {QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning},
  author = {Hossein Rajabzadeh and Mojtaba Valipour and Tianshu Zhu and Marzieh Tahaei and Hyock Ju Kwon and Ali Ghodsi and Boxing Chen and Mehdi Rezagholizadeh},
  journal= {arXiv preprint arXiv:2402.10462},
  year   = {2025}
}

Comments

Best Paper Award AAAI EIW Workshop

R2 v1 2026-06-28T14:50:22.447Z