English

Contrastive learning of T cell receptor representations

Biomolecules 2024-10-11 v2 Artificial Intelligence Machine Learning

Abstract

Computational prediction of the interaction of T cell receptors (TCRs) and their ligands is a grand challenge in immunology. Despite advances in high-throughput assays, specificity-labelled TCR data remains sparse. In other domains, the pre-training of language models on unlabelled data has been successfully used to address data bottlenecks. However, it is unclear how to best pre-train protein language models for TCR specificity prediction. Here we introduce a TCR language model called SCEPTR (Simple Contrastive Embedding of the Primary sequence of T cell Receptors), capable of data-efficient transfer learning. Through our model, we introduce a novel pre-training strategy combining autocontrastive learning and masked-language modelling, which enables SCEPTR to achieve its state-of-the-art performance. In contrast, existing protein language models and a variant of SCEPTR pre-trained without autocontrastive learning are outperformed by sequence alignment-based methods. We anticipate that contrastive learning will be a useful paradigm to decode the rules of TCR specificity.

Keywords

Cite

@article{arxiv.2406.06397,
  title  = {Contrastive learning of T cell receptor representations},
  author = {Yuta Nagano and Andrew Pyo and Martina Milighetti and James Henderson and John Shawe-Taylor and Benny Chain and Andreas Tiffeau-Mayer},
  journal= {arXiv preprint arXiv:2406.06397},
  year   = {2024}
}

Comments

25 pages, 23 figures; additional analyses and improvements to existing figures

R2 v1 2026-06-28T16:59:49.494Z