English

Procedural generation of meta-reinforcement learning tasks

Machine Learning 2023-12-12 v2 Artificial Intelligence

Abstract

Open-endedness stands to benefit from the ability to generate an infinite variety of diverse, challenging environments. One particularly interesting type of challenge is meta-learning ("learning-to-learn"), a hallmark of intelligent behavior. However, the number of meta-learning environments in the literature is limited. Here we describe a parametrized space for simple meta-reinforcement learning (meta-RL) tasks with arbitrary stimuli. The parametrization allows us to randomly generate an arbitrary number of novel simple meta-learning tasks. The parametrization is expressive enough to include many well-known meta-RL tasks, such as bandit problems, the Harlow task, T-mazes, the Daw two-step task and others. Simple extensions allow it to capture tasks based on two-dimensional topological spaces, such as full mazes or find-the-spot domains. We describe a number of randomly generated meta-RL domains of varying complexity and discuss potential issues arising from random generation.

Keywords

Cite

@article{arxiv.2302.05583,
  title  = {Procedural generation of meta-reinforcement learning tasks},
  author = {Thomas Miconi},
  journal= {arXiv preprint arXiv:2302.05583},
  year   = {2023}
}

Comments

Agent Learning in Open-Endedness (ALOE) Workshop at NeurIPS 2023

R2 v1 2026-06-28T08:37:33.341Z