Diversity Policy Gradient for Sample Efficient Quality-Diversity Optimization

Thomas Pierrot; Valentin Macé; Félix Chalumeau; Arthur Flajolet; Geoffrey Cideron; Karim Beguir; Antoine Cully; Olivier Sigaud; Nicolas Perrin-Gilbert

doi:10.1145/3512290.3528845

Diversity Policy Gradient for Sample Efficient Quality-Diversity Optimization

Artificial Intelligence 2022-06-01 v5 Machine Learning

Authors: Thomas Pierrot , Valentin Macé , Félix Chalumeau , Arthur Flajolet , Geoffrey Cideron , Karim Beguir , Antoine Cully , Olivier Sigaud , Nicolas Perrin-Gilbert

View on arXiv ↗ PDF ↗ DOI ↗

Abstract

A fascinating aspect of nature lies in its ability to produce a large and diverse collection of organisms that are all high-performing in their niche. By contrast, most AI algorithms focus on finding a single efficient solution to a given problem. Aiming for diversity in addition to performance is a convenient way to deal with the exploration-exploitation trade-off that plays a central role in learning. It also allows for increased robustness when the returned collection contains several working solutions to the considered problem, making it well-suited for real applications such as robotics. Quality-Diversity (QD) methods are evolutionary algorithms designed for this purpose. This paper proposes a novel algorithm, QDPG, which combines the strength of Policy Gradient algorithms and Quality Diversity approaches to produce a collection of diverse and high-performing neural policies in continuous control environments. The main contribution of this work is the introduction of a Diversity Policy Gradient (DPG) that exploits information at the time-step level to drive policies towards more diversity in a sample-efficient manner. Specifically, QDPG selects neural controllers from a MAP-Elites grid and uses two gradient-based mutation operators to improve both quality and diversity. Our results demonstrate that QDPG is significantly more sample-efficient than its evolutionary competitors.

Keywords

quality diversity algorithms evolutionary algorithm genetic algorithm

Cite

@article{arxiv.2006.08505,
  title  = {Diversity Policy Gradient for Sample Efficient Quality-Diversity Optimization},
  author = {Thomas Pierrot and Valentin Macé and Félix Chalumeau and Arthur Flajolet and Geoffrey Cideron and Karim Beguir and Antoine Cully and Olivier Sigaud and Nicolas Perrin-Gilbert},
  journal= {arXiv preprint arXiv:2006.08505},
  year   = {2022}
}

Comments

Add several baselines (Policy Gradient assisted MAP Elites, DIAYN, AGAC) Change writing to take the point of view of the evo community Change style, writing, explanation, figures

Diversity Policy Gradient for Sample Efficient Quality-Diversity Optimization

Abstract

Keywords

Cite

Comments

Related papers