English

Exploring Protein Language Model Architecture-Induced Biases for Antibody Comprehension

Machine Learning 2025-12-11 v1

Abstract

Recent advances in protein language models (PLMs) have demonstrated remarkable capabilities in understanding protein sequences. However, the extent to which different model architectures capture antibody-specific biological properties remains unexplored. In this work, we systematically investigate how architectural choices in PLMs influence their ability to comprehend antibody sequence characteristics and functions. We evaluate three state-of-the-art PLMs-AntiBERTa, BioBERT, and ESM2--against a general-purpose language model (GPT-2) baseline on antibody target specificity prediction tasks. Our results demonstrate that while all PLMs achieve high classification accuracy, they exhibit distinct biases in capturing biological features such as V gene usage, somatic hypermutation patterns, and isotype information. Through attention attribution analysis, we show that antibody-specific models like AntiBERTa naturally learn to focus on complementarity-determining regions (CDRs), while general protein models benefit significantly from explicit CDR-focused training strategies. These findings provide insights into the relationship between model architecture and biological feature extraction, offering valuable guidance for future PLM development in computational antibody design.

Keywords

Cite

@article{arxiv.2512.09894,
  title  = {Exploring Protein Language Model Architecture-Induced Biases for Antibody Comprehension},
  author = {Mengren and Liu and Yixiang Zhang and Yiming and Zhang},
  journal= {arXiv preprint arXiv:2512.09894},
  year   = {2025}
}
R2 v1 2026-07-01T08:19:14.635Z