中文

利用强化学习在可合成化学空间中导航

机器学习 2020-05-21 v2 人工智能

摘要

过去十年中,用于从头药物设计的机器学习领域取得了显著进展,尤其在深度生成模型方面。然而,当前的生成式方法面临一大挑战:它们无法确保所提出的分子结构可被实际合成,也不提供所提出小分子的合路线,从而严重限制了其实用性。本工作中,我们提出了一种由强化学习(RL)驱动、用于从头药物设计的新型前向合成框架——前向合成策略梯度(PGFS),其通过将合成可及性概念直接嵌入从头药物设计系统来应对该挑战。在此设定下,智能体在迭代式虚拟多步合成过程的每一步,对市售小分子构建块施加有效化学反应,从而学会在庞大的可合成化学空间中导航。所提出的药物发现环境因具有含层次动作的大状态空间与高维连续动作空间,为 RL 算法提供了极具挑战的测试平台。PGFS 在生成具有高 QED 与惩罚性 clogP 的结构方面达到了 SOTA 性能。此外,我们在涉及三个 HIV 靶点的计算机概念验证中验证了 PGFS。最后,我们阐述了本研究所概念化的端到端训练如何代表一种重要范式,可从根本上扩展可合成化学空间并自动化药物发现流程。

关键词

引用

@article{arxiv.2004.12485,
  title  = {Learning To Navigate The Synthetically Accessible Chemical Space Using Reinforcement Learning},
  author = {Sai Krishna Gottipati and Boris Sattarov and Sufeng Niu and Yashaswi Pathak and Haoran Wei and Shengchao Liu and Karam M. J. Thomas and Simon Blackburn and Connor W. Coley and Jian Tang and Sarath Chandar and Yoshua Bengio},
  journal= {arXiv preprint arXiv:2004.12485},
  year   = {2020}
}

备注

added the statistics of top-100 compounds used logP metric with scaled components added values of the initial reactants to the box plots some values in tables are recalculated due to the inconsistent environments on different machines. corresponding benchmarks were rerun with the requirements on github. no significant changes in the results. corrected figures in the Appendix