English

Gated Multimodal Units for Information Fusion

Machine Learning 2017-02-08 v1 Machine Learning

Abstract

This paper presents a novel model for multimodal learning based on gated neural networks. The Gated Multimodal Unit (GMU) model is intended to be used as an internal unit in a neural network architecture whose purpose is to find an intermediate representation based on a combination of data from different modalities. The GMU learns to decide how modalities influence the activation of the unit using multiplicative gates. It was evaluated on a multilabel scenario for genre classification of movies using the plot and the poster. The GMU improved the macro f-score performance of single-modality approaches and outperformed other fusion strategies, including mixture of experts models. Along with this work, the MM-IMDb dataset is released which, to the best of our knowledge, is the largest publicly available multimodal dataset for genre prediction on movies.

Keywords

Cite

@article{arxiv.1702.01992,
  title  = {Gated Multimodal Units for Information Fusion},
  author = {John Arevalo and Thamar Solorio and Manuel Montes-y-Gómez and Fabio A. González},
  journal= {arXiv preprint arXiv:1702.01992},
  year   = {2017}
}
R2 v1 2026-06-22T18:11:30.526Z