English

Combining Residual Networks with LSTMs for Lipreading

Computer Vision and Pattern Recognition 2017-09-11 v4

Abstract

We propose an end-to-end deep learning architecture for word-level visual speech recognition. The system is a combination of spatiotemporal convolutional, residual and bidirectional Long Short-Term Memory networks. We train and evaluate it on the Lipreading In-The-Wild benchmark, a challenging database of 500-size target-words consisting of 1.28sec video excerpts from BBC TV broadcasts. The proposed network attains word accuracy equal to 83.0, yielding 6.8 absolute improvement over the current state-of-the-art, without using information about word boundaries during training or testing.

Keywords

Cite

@article{arxiv.1703.04105,
  title  = {Combining Residual Networks with LSTMs for Lipreading},
  author = {Themos Stafylakis and Georgios Tzimiropoulos},
  journal= {arXiv preprint arXiv:1703.04105},
  year   = {2017}
}

Comments

Submitted to Interspeech 2017

R2 v1 2026-06-22T18:43:26.370Z