连续时间连续空间稳态强化学习 (CTCS-HRRL):迈向生物自自主智能体
人工智能
2024-01-18 v1 机器学习
摘要
稳态是生物体维持内部平衡的生物学过程。先前的研究表明,稳态是一种习得行为。最近提出的稳态调节强化学习 (HRRL) 框架试图通过连接驱力减少理论 (Drive Reduction Theory) 和强化学习来解释这种习得的稳态行为。这种联系已在离散时空得到证明,但在连续时空中尚未得到验证。在本工作中,我们将 HRRL 框架推进到连续时空环境,并验证了 CTCS-HRRL(连续时间连续空间 HRRL)框架。我们通过设计一个模仿真实世界生物代理稳态机制的模型来实现这一目标。该模型使用 Hamilton-Jacobian-Bellman 方程,以及基于神经网络和强化学习的函数逼近。通过基于模拟的实验,我们展示了该模型的有效性,并揭示了与代理在持续变化的内部状态环境中动态选择有利于稳态的策略的能力相关的证据。我们的实验结果表明,代理在 CTCS 环境中学习了稳态行为,使得 CTCS-HRRL 成为模拟动物动力学和决策的一个有前景的框架。
引用
@article{arxiv.2401.08999,
title = {Continuous Time Continuous Space Homeostatic Reinforcement Learning (CTCS-HRRL) : Towards Biological Self-Autonomous Agent},
author = {Hugo Laurencon and Yesoda Bhargava and Riddhi Zantye and Charbel-Raphaël Ségerie and Johann Lussange and Veeky Baths and Boris Gutkin},
journal= {arXiv preprint arXiv:2401.08999},
year = {2024}
}
备注
This work is a result of the ongoing collaboration between Cognitive Neuroscience Lab, BITS Pilani K K Birla Goa Campus and Ecole Normale Superieure, Paris France. This work is jointly supervised by Prof. Boris Gutkin and Prof. Veeky Baths. arXiv admin note: substantial text overlap with arXiv:2109.06580