English

HoME: a Household Multimodal Environment

Artificial Intelligence 2017-11-30 v1 Computation and Language Computer Vision and Pattern Recognition Robotics Sound Audio and Speech Processing

Abstract

We introduce HoME: a Household Multimodal Environment for artificial agents to learn from vision, audio, semantics, physics, and interaction with objects and other agents, all within a realistic context. HoME integrates over 45,000 diverse 3D house layouts based on the SUNCG dataset, a scale which may facilitate learning, generalization, and transfer. HoME is an open-source, OpenAI Gym-compatible platform extensible to tasks in reinforcement learning, language grounding, sound-based navigation, robotics, multi-agent learning, and more. We hope HoME better enables artificial agents to learn as humans do: in an interactive, multimodal, and richly contextualized setting.

Keywords

Cite

@article{arxiv.1711.11017,
  title  = {HoME: a Household Multimodal Environment},
  author = {Simon Brodeur and Ethan Perez and Ankesh Anand and Florian Golemo and Luca Celotti and Florian Strub and Jean Rouat and Hugo Larochelle and Aaron Courville},
  journal= {arXiv preprint arXiv:1711.11017},
  year   = {2017}
}

Comments

Presented at NIPS 2017's Visually-Grounded Interaction and Language Workshop

R2 v1 2026-06-22T23:01:21.394Z