English

Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs

Computation and Language 2025-08-26 v1

Abstract

Automatic Speech Recognition has advanced with self-supervised learning, enabling feature extraction directly from raw audio. In Wav2Vec, a CNN first transforms audio into feature vectors before the transformer processes them. This study examines CNN-extracted information for monophthong vowels using the TIMIT corpus. We compare MFCCs, MFCCs with formants, and CNN activations by training SVM classifiers for front-back vowel identification, assessing their classification accuracy to evaluate phonetic representation.

Keywords

Cite

@article{arxiv.2508.17914,
  title  = {Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs},
  author = {Domenico De Cristofaro and Vincenzo Norman Vitale and Alessandro Vietti},
  journal= {arXiv preprint arXiv:2508.17914},
  year   = {2025}
}
R2 v1 2026-07-01T05:04:25.908Z