English

Preference optimization of protein language models as a multi-objective binder design paradigm

Biological Physics 2024-03-08 v1 Artificial Intelligence Computational Engineering, Finance, and Science Biomolecules

Abstract

We present a multi-objective binder design paradigm based on instruction fine-tuning and direct preference optimization (DPO) of autoregressive protein language models (pLMs). Multiple design objectives are encoded in the language model through direct optimization on expert curated preference sequence datasets comprising preferred and dispreferred distributions. We show the proposed alignment strategy enables ProtGPT2 to effectively design binders conditioned on specified receptors and a drug developability criterion. Generated binder samples demonstrate median isoelectric point (pI) improvements by 17%60%17\%-60\%.

Keywords

Cite

@article{arxiv.2403.04187,
  title  = {Preference optimization of protein language models as a multi-objective binder design paradigm},
  author = {Pouria Mistani and Venkatesh Mysore},
  journal= {arXiv preprint arXiv:2403.04187},
  year   = {2024}
}

Comments

Published at the GEM workshop, ICLR 2024. Generative and Experimental Perspectives for Biomolecular Design (https://www.gembio.ai/)

R2 v1 2026-06-28T15:11:47.041Z