English

Speaker conditioned acoustic-to-articulatory inversion using x-vectors

Audio and Speech Processing 2020-06-23 v1

Abstract

Speech production involves the movement of various articulators, including tongue, jaw, and lips. Estimating the movement of the articulators from the acoustics of speech is known as acoustic-to-articulatory inversion (AAI). Recently, it has been shown that instead of training AAI in a speaker specific manner, pooling the acoustic-articulatory data from multiple speakers is beneficial. Further, additional conditioning with speaker specific information by one-hot encoding at the input of AAI along with acoustic features benefits the AAI performance in a closed-set speaker train and test condition. In this work, we carry out an experimental study on the benefit of using x-vectors for providing speaker specific information to condition AAI. Experiments with 30 speakers have shown that the AAI performance benefits from the use of x-vectors in a closed set seen speaker condition. Further, x-vectors also generalizes well for unseen speaker evaluation.

Keywords

Cite

@article{arxiv.2006.11536,
  title  = {Speaker conditioned acoustic-to-articulatory inversion using x-vectors},
  author = {Aravind Illa and Prasanta Kumar Ghosh},
  journal= {arXiv preprint arXiv:2006.11536},
  year   = {2020}
}
R2 v1 2026-06-23T16:29:04.535Z