English

Deep Multimodal Semantic Embeddings for Speech and Images

Computer Vision and Pattern Recognition 2015-11-13 v1 Artificial Intelligence Computation and Language

Abstract

In this paper, we present a model which takes as input a corpus of images with relevant spoken captions and finds a correspondence between the two modalities. We employ a pair of convolutional neural networks to model visual objects and speech signals at the word level, and tie the networks together with an embedding and alignment model which learns a joint semantic space over both modalities. We evaluate our model using image search and annotation tasks on the Flickr8k dataset, which we augmented by collecting a corpus of 40,000 spoken captions using Amazon Mechanical Turk.

Keywords

Cite

@article{arxiv.1511.03690,
  title  = {Deep Multimodal Semantic Embeddings for Speech and Images},
  author = {David Harwath and James Glass},
  journal= {arXiv preprint arXiv:1511.03690},
  year   = {2015}
}
R2 v1 2026-06-22T11:43:01.965Z