English

Fine-tuning Protein Language Models with Deep Mutational Scanning improves Variant Effect Prediction

Genomics 2024-05-14 v1 Machine Learning

Abstract

Protein Language Models (PLMs) have emerged as performant and scalable tools for predicting the functional impact and clinical significance of protein-coding variants, but they still lag experimental accuracy. Here, we present a novel fine-tuning approach to improve the performance of PLMs with experimental maps of variant effects from Deep Mutational Scanning (DMS) assays using a Normalised Log-odds Ratio (NLR) head. We find consistent improvements in a held-out protein test set, and on independent DMS and clinical variant annotation benchmarks from ProteinGym and ClinVar. These findings demonstrate that DMS is a promising source of sequence diversity and supervised training data for improving the performance of PLMs for variant effect prediction.

Keywords

Cite

@article{arxiv.2405.06729,
  title  = {Fine-tuning Protein Language Models with Deep Mutational Scanning improves Variant Effect Prediction},
  author = {Aleix Lafita and Ferran Gonzalez and Mahmoud Hossam and Paul Smyth and Jacob Deasy and Ari Allyn-Feuer and Daniel Seaton and Stephen Young},
  journal= {arXiv preprint arXiv:2405.06729},
  year   = {2024}
}

Comments

Machine Learning for Genomics Explorations workshop at ICLR 2024

R2 v1 2026-06-28T16:23:39.361Z