English

ARTI-6: Towards Six-dimensional Articulatory Speech Encoding

Audio and Speech Processing 2026-01-27 v2 Artificial Intelligence Computation and Language

Abstract

We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, tongue root, and larynx. ARTI-6 consists of three components: (1) a six-dimensional articulatory feature set representing key regions of the vocal tract; (2) an articulatory inversion model, which predicts articulatory features from speech acoustics leveraging speech foundation models, achieving a prediction correlation of 0.87; and (3) an articulatory synthesis model, which reconstructs intelligible speech directly from articulatory features, showing that even a low-dimensional representation can generate natural-sounding speech. Together, ARTI-6 provides an interpretable, computationally efficient, and physiologically grounded framework for advancing articulatory inversion, synthesis, and broader speech technology applications. The source code and speech samples are publicly available.

Keywords

Cite

@article{arxiv.2509.21447,
  title  = {ARTI-6: Towards Six-dimensional Articulatory Speech Encoding},
  author = {Jihwan Lee and Sean Foley and Thanathai Lertpetchpun and Kevin Huang and Yoonjeong Lee and Tiantian Feng and Louis Goldstein and Dani Byrd and Shrikanth Narayanan},
  journal= {arXiv preprint arXiv:2509.21447},
  year   = {2026}
}

Comments

Accepted for ICASSP 2026

R2 v1 2026-07-01T05:56:51.978Z