English

Is Sequence Information All You Need for Bayesian Optimization of Antibodies?

Machine Learning 2025-09-30 v1 Quantitative Methods

Abstract

Bayesian optimization is a natural candidate for the engineering of antibody therapeutic properties, which is often iterative and expensive. However, finding the optimal choice of surrogate model for optimization over the highly structured antibody space is difficult, and may differ depending on the property being optimized. Moreover, to the best of our knowledge, no prior works have attempted to incorporate structural information into antibody Bayesian optimization. In this work, we explore different approaches to incorporating structural information into Bayesian optimization, and compare them to a variety of sequence-only approaches on two different antibody properties, binding affinity and stability. In addition, we propose the use of a protein language model-based ``soft constraint,'' which helps guide the optimization to promising regions of the space. We find that certain types of structural information improve data efficiency in early optimization rounds for stability, but have equivalent peak performance. Moreover, when incorporating the protein language model soft constraint we find that the data efficiency gap is diminished for affinity and eliminated for stability, resulting in sequence-only methods that match the performance of structure-based methods, raising questions about the necessity of structure in Bayesian optimization for antibodies.

Keywords

Cite

@article{arxiv.2509.24933,
  title  = {Is Sequence Information All You Need for Bayesian Optimization of Antibodies?},
  author = {Sebastian W. Ober and Calvin McCarter and Aniruddh Raghu and Yucen Lily Li and Alan N. Amin and Andrew Gordon Wilson and Hunter Elliott},
  journal= {arXiv preprint arXiv:2509.24933},
  year   = {2025}
}

Comments

Accepted into the AI for Science Workshop, NeurIPS 2025

R2 v1 2026-07-01T06:04:51.190Z