English

One-Shot Learning from a Demonstration with Hierarchical Latent Language

Computation and Language 2022-03-10 v1

Abstract

Humans have the capability, aided by the expressive compositionality of their language, to learn quickly by demonstration. They are able to describe unseen task-performing procedures and generalize their execution to other contexts. In this work, we introduce DescribeWorld, an environment designed to test this sort of generalization skill in grounded agents, where tasks are linguistically and procedurally composed of elementary concepts. The agent observes a single task demonstration in a Minecraft-like grid world, and is then asked to carry out the same task in a new map. To enable such a level of generalization, we propose a neural agent infused with hierarchical latent language--both at the level of task inference and subtask planning. Our agent first generates a textual description of the demonstrated unseen task, then leverages this description to replicate it. Through multiple evaluation scenarios and a suite of generalization tests, we find that agents that perform text-based inference are better equipped for the challenge under a random split of tasks.

Keywords

Cite

@article{arxiv.2203.04806,
  title  = {One-Shot Learning from a Demonstration with Hierarchical Latent Language},
  author = {Nathaniel Weir and Xingdi Yuan and Marc-Alexandre Côté and Matthew Hausknecht and Romain Laroche and Ida Momennejad and Harm Van Seijen and Benjamin Van Durme},
  journal= {arXiv preprint arXiv:2203.04806},
  year   = {2022}
}
R2 v1 2026-06-24T10:07:29.459Z