中文

通过 In-context Learning 和 Test-time Training 进行 Few-shot 蛋白质适应性预测

生物大分子 2025-12-03 v1 机器学习

摘要

准确预测在实验数据有限的情况下的蛋白质适应性是蛋白质工程中的持久挑战。我们提出了 PRIMO (PRotein In-context Mutation Oracle),一个基于 transformer 的框架,利用 in-context learning 和 test-time training 在不依赖大规模任务特定数据集的情况下快速适应新蛋白质和 assay。通过将序列信息、辅助零时预测以及来自多个 assay 的稀疏实验标签编码为统一的 token 集合于预训练的掩码语言模型范式中,PRIMO 学习通过基于偏好的损失函数来优先挑选有前景的变体。在 diverse protein families 和 properties 方面,包括 substitution 和 indel 突变,PRIMO 的性能优于零时和 fully supervised baseline。本工作强调了将大规模预训练与高效 test-time 适应相结合的力量,以应对数据收集昂贵且标签稀缺的挑战性蛋白质设计任务。

关键词

引用

@article{arxiv.2512.02315,
  title  = {Few-shot Protein Fitness Prediction via In-context Learning and Test-time Training},
  author = {Felix Teufel and Aaron W. Kollasch and Yining Huang and Ole Winther and Kevin K. Yang and Pascal Notin and Debora S. Marks},
  journal= {arXiv preprint arXiv:2512.02315},
  year   = {2025}
}

备注

AI for Science Workshop (NeurIPS 2025)