Self-supervised model pre-training has recently garnered significant interest, but relatively few efforts have explored using additional resources in fine-tuning these models. We demonstrate how universal phoneset acoustic models can leverage cross-lingual supervision to improve transfer of pretrained self-supervised representations to new languages. We also show how target-language text can be used to enable and improve fine-tuning with the lattice-free maximum mutual information (LF-MMI) objective. In three low-resource languages these techniques greatly improved few-shot learning performance.
@article{arxiv.2110.04863,
title = {Injecting Text and Cross-lingual Supervision in Few-shot Learning from Self-Supervised Models},
author = {Matthew Wiesner and Desh Raj and Sanjeev Khudanpur},
journal= {arXiv preprint arXiv:2110.04863},
year = {2021}
}
Comments
\c{opyright} 2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works