English

Can DNNs Learn to Lipread Full Sentences?

Image and Video Processing 2018-05-31 v1 Computer Vision and Pattern Recognition Audio and Speech Processing

Abstract

Finding visual features and suitable models for lipreading tasks that are more complex than a well-constrained vocabulary has proven challenging. This paper explores state-of-the-art Deep Neural Network architectures for lipreading based on a Sequence to Sequence Recurrent Neural Network. We report results for both hand-crafted and 2D/3D Convolutional Neural Network visual front-ends, online monotonic attention, and a joint Connectionist Temporal Classification-Sequence-to-Sequence loss. The system is evaluated on the publicly available TCD-TIMIT dataset, with 59 speakers and a vocabulary of over 6000 words. Results show a major improvement on a Hidden Markov Model framework. A fuller analysis of performance across visemes demonstrates that the network is not only learning the language model, but actually learning to lipread.

Keywords

Cite

@article{arxiv.1805.11685,
  title  = {Can DNNs Learn to Lipread Full Sentences?},
  author = {George Sterpu and Christian Saam and Naomi Harte},
  journal= {arXiv preprint arXiv:1805.11685},
  year   = {2018}
}

Comments

Accepted at the 2018 IEEE International Conference on Image Processing (ICIP 2018)

R2 v1 2026-06-23T02:12:34.435Z