English

Low-rank Adaptation Method for Wav2vec2-based Fake Audio Detection

Sound 2023-06-12 v1 Computation and Language Audio and Speech Processing

Abstract

Self-supervised speech models are a rapidly developing research topic in fake audio detection. Many pre-trained models can serve as feature extractors, learning richer and higher-level speech features. However,when fine-tuning pre-trained models, there is often a challenge of excessively long training times and high memory consumption, and complete fine-tuning is also very expensive. To alleviate this problem, we apply low-rank adaptation(LoRA) to the wav2vec2 model, freezing the pre-trained model weights and injecting a trainable rank-decomposition matrix into each layer of the transformer architecture, greatly reducing the number of trainable parameters for downstream tasks. Compared with fine-tuning with Adam on the wav2vec2 model containing 317M training parameters, LoRA achieved similar performance by reducing the number of trainable parameters by 198 times.

Keywords

Cite

@article{arxiv.2306.05617,
  title  = {Low-rank Adaptation Method for Wav2vec2-based Fake Audio Detection},
  author = {Chenglong Wang and Jiangyan Yi and Xiaohui Zhang and Jianhua Tao and Le Xu and Ruibo Fu},
  journal= {arXiv preprint arXiv:2306.05617},
  year   = {2023}
}

Comments

6pages

R2 v1 2026-06-28T11:00:38.634Z