English

Deep Neural Convolutive Matrix Factorization for Articulatory Representation Decomposition

Audio and Speech Processing 2022-06-22 v3 Artificial Intelligence Signal Processing

Abstract

Most of the research on data-driven speech representation learning has focused on raw audios in an end-to-end manner, paying little attention to their internal phonological or gestural structure. This work, investigating the speech representations derived from articulatory kinematics signals, uses a neural implementation of convolutive sparse matrix factorization to decompose the articulatory data into interpretable gestures and gestural scores. By applying sparse constraints, the gestural scores leverage the discrete combinatorial properties of phonological gestures. Phoneme recognition experiments were additionally performed to show that gestural scores indeed code phonological information successfully. The proposed work thus makes a bridge between articulatory phonology and deep neural networks to leverage informative, intelligible, interpretable,and efficient speech representations.

Keywords

Cite

@article{arxiv.2204.00465,
  title  = {Deep Neural Convolutive Matrix Factorization for Articulatory Representation Decomposition},
  author = {Jiachen Lian and Alan W Black and Louis Goldstein and Gopala Krishna Anumanchipalli},
  journal= {arXiv preprint arXiv:2204.00465},
  year   = {2022}
}

Comments

Accepted to 2022 Interspeech. Code is publicly available at https://github.com/Berkeley-Speech-Group/ema_gesture

R2 v1 2026-06-24T10:34:45.456Z