Exploring Phoneme-Level Speech Representations for End-to-End Speech Translation
Computation and Language
2019-06-05 v1 Sound
Audio and Speech Processing
Abstract
Previous work on end-to-end translation from speech has primarily used frame-level features as speech representations, which creates longer, sparser sequences than text. We show that a naive method to create compressed phoneme-like speech representations is far more effective and efficient for translation than traditional frame-level speech features. Specifically, we generate phoneme labels for speech frames and average consecutive frames with the same label to create shorter, higher-level source sequences for translation. We see improvements of up to 5 BLEU on both our high and low resource language pairs, with a reduction in training time of 60%. Our improvements hold across multiple data sizes and two language pairs.
Cite
@article{arxiv.1906.01199,
title = {Exploring Phoneme-Level Speech Representations for End-to-End Speech Translation},
author = {Elizabeth Salesky and Matthias Sperber and Alan W Black},
journal= {arXiv preprint arXiv:1906.01199},
year = {2019}
}
Comments
Accepted to ACL 2019