中文

将深度强化学习应用于 HP 模型进行蛋白质结构预测

机器学习 2023-03-14 v2 生物大分子

摘要

计算生物物理学中的一个核心问题是蛋白质结构预测,即寻找给定氨基酸序列的最优折叠。该问题已在一个经典的抽象模型——HP 模型中被研究,其中蛋白质被建模为晶格上由 H(疏水)和 P(极性)氨基酸组成的序列。目标是找到使 H-H 接触最大化的构象。已知即使在这种简化设定下,该问题仍是难解的(NP-hard)。在本工作中,我们将深度强化学习(DRL)应用于二维 HP 模型。我们能够为长度从 20 到 50 的基准 HP 序列获得已知最优能量的构象。我们的 DRL 基于深度 Q 网络(DQN)。我们发现,基于长短期记忆(LSTM)架构的 DQN 极大增强了 RL 的学习能力并显著改进了搜索过程。DRL 能够高效地对状态空间进行采样,而无需人工启发式方法。实验上我们表明,它每次试验都能找到多个不同的最优已知解。本研究展示了深度强化学习在蛋白质折叠 HP 模型中的有效性。

关键词

引用

@article{arxiv.2211.14939,
  title  = {Applying Deep Reinforcement Learning to the HP Model for Protein Structure Prediction},
  author = {Kaiyuan Yang and Houjing Huang and Olafs Vandans and Adithya Murali and Fujia Tian and Roland H. C. Yap and Liang Dai},
  journal= {arXiv preprint arXiv:2211.14939},
  year   = {2023}
}

备注

Published at Physica A: Statistical Mechanics and its Applications, available online 7 December 2022. Extended abstract accepted by the Machine Learning and the Physical Sciences workshop, NeurIPS 2022