From Optimal Policies to Individual Differences: Rethinking Reinforcement Learning for Biology
Neural and Evolutionary Computing
2026-07-17 v1
Abstract
Reinforcement learning (RL) is primarily known as a computational method for optimizing control tasks, but it is increasingly used to explain biological behavior. While RL successfully captures key aspects of biology, a major gap remains: between-agent behavioral variability. Consistent individual differences naturally permeate biological populations, yet RL models typically present only the single best individual or the population average. Addressing this gap requires moving beyond current practices to generate behavioral diversity using biologically plausible mechanisms. Here, we examine approaches from various subfields of RL and outline potential paths forward to close the gap between biology and simulation.
Keywords
Cite
@article{arxiv.2607.16542,
title = {From Optimal Policies to Individual Differences: Rethinking Reinforcement Learning for Biology},
author = {Patrick Govoni and Palina Bartashevich and Clémence Bergerot and Valerii Chirkov and Valentin Lecheval and Pawel Romanczuk},
journal= {arXiv preprint arXiv:2607.16542},
year = {2026}
}