English

Explain Images with Multimodal Recurrent Neural Networks

Computer Vision and Pattern Recognition 2014-10-07 v1 Computation and Language Machine Learning

Abstract

In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel sentence descriptions to explain the content of images. It directly models the probability distribution of generating a word given previous words and the image. Image descriptions are generated by sampling from this distribution. The model consists of two sub-networks: a deep recurrent neural network for sentences and a deep convolutional network for images. These two sub-networks interact with each other in a multimodal layer to form the whole m-RNN model. The effectiveness of our model is validated on three benchmark datasets (IAPR TC-12, Flickr 8K, and Flickr 30K). Our model outperforms the state-of-the-art generative method. In addition, the m-RNN model can be applied to retrieval tasks for retrieving images or sentences, and achieves significant performance improvement over the state-of-the-art methods which directly optimize the ranking objective function for retrieval.

Keywords

Cite

@article{arxiv.1410.1090,
  title  = {Explain Images with Multimodal Recurrent Neural Networks},
  author = {Junhua Mao and Wei Xu and Yi Yang and Jiang Wang and Alan L. Yuille},
  journal= {arXiv preprint arXiv:1410.1090},
  year   = {2014}
}
R2 v1 2026-06-22T06:13:11.322Z