English

MFAS: Multimodal Fusion Architecture Search

Machine Learning 2019-03-18 v1 Computer Vision and Pattern Recognition Neural and Evolutionary Computing

Abstract

We tackle the problem of finding good architectures for multimodal classification problems. We propose a novel and generic search space that spans a large number of possible fusion architectures. In order to find an optimal architecture for a given dataset in the proposed search space, we leverage an efficient sequential model-based exploration approach that is tailored for the problem. We demonstrate the value of posing multimodal fusion as a neural architecture search problem by extensive experimentation on a toy dataset and two other real multimodal datasets. We discover fusion architectures that exhibit state-of-the-art performance for problems with different domain and dataset size, including the NTU RGB+D dataset, the largest multi-modal action recognition dataset available.

Keywords

Cite

@article{arxiv.1903.06496,
  title  = {MFAS: Multimodal Fusion Architecture Search},
  author = {Juan-Manuel Pérez-Rúa and Valentin Vielzeuf and Stéphane Pateux and Moez Baccouche and Frédéric Jurie},
  journal= {arXiv preprint arXiv:1903.06496},
  year   = {2019}
}

Comments

CVPR 2019, Jun 2019, Long Beach, United States http://cvpr2019.thecvf.com/

R2 v1 2026-06-23T08:09:17.938Z