English

Improving Speaker-Independent Lipreading with Domain-Adversarial Training

Computer Vision and Pattern Recognition 2017-08-07 v1 Computation and Language

Abstract

We present a Lipreading system, i.e. a speech recognition system using only visual features, which uses domain-adversarial training for speaker independence. Domain-adversarial training is integrated into the optimization of a lipreader based on a stack of feedforward and LSTM (Long Short-Term Memory) recurrent neural networks, yielding an end-to-end trainable system which only requires a very small number of frames of untranscribed target data to substantially improve the recognition accuracy on the target speaker. On pairs of different source and target speakers, we achieve a relative accuracy improvement of around 40% with only 15 to 20 seconds of untranscribed target speech data. On multi-speaker training setups, the accuracy improvements are smaller but still substantial.

Keywords

Cite

@article{arxiv.1708.01565,
  title  = {Improving Speaker-Independent Lipreading with Domain-Adversarial Training},
  author = {Michael Wand and Juergen Schmidhuber},
  journal= {arXiv preprint arXiv:1708.01565},
  year   = {2017}
}

Comments

Accepted at Interspeech 2017

R2 v1 2026-06-22T21:07:11.184Z