English

An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model

Computer Vision and Pattern Recognition 2024-05-21 v2

Abstract

Recent studies applied Parameter Efficient Fine-Tuning techniques (PEFTs) to efficiently narrow the performance gap between pre-training and downstream. There are two important factors for various PEFTs, namely, the accessible data size and fine-tunable parameter size. A natural expectation for PEFTs is that the performance of various PEFTs is positively related to the data size and fine-tunable parameter size. However, according to the evaluation of five PEFTs on two downstream vision-language (VL) tasks, we find that such an intuition holds only if the downstream data and task are not consistent with pre-training. For downstream fine-tuning consistent with pre-training, data size no longer affects the performance, while the influence of fine-tunable parameter size is not monotonous. We believe such an observation could guide the choice of training strategy for various PEFTs.

Keywords

Cite

@article{arxiv.2403.08433,
  title  = {An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model},
  author = {Yuxin Tian and Mouxing Yang and Yunfan Li and Dayiheng Liu and Xingzhang Ren and Xi Peng and Jiancheng Lv},
  journal= {arXiv preprint arXiv:2403.08433},
  year   = {2024}
}

Comments

Accepted by ICME2024

R2 v1 2026-06-28T15:18:34.689Z