Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
Computation and Language
2025-08-26 v1
Abstract
Automatic Speech Recognition has advanced with self-supervised learning, enabling feature extraction directly from raw audio. In Wav2Vec, a CNN first transforms audio into feature vectors before the transformer processes them. This study examines CNN-extracted information for monophthong vowels using the TIMIT corpus. We compare MFCCs, MFCCs with formants, and CNN activations by training SVM classifiers for front-back vowel identification, assessing their classification accuracy to evaluate phonetic representation.
Keywords
Cite
@article{arxiv.2508.17914,
title = {Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs},
author = {Domenico De Cristofaro and Vincenzo Norman Vitale and Alessandro Vietti},
journal= {arXiv preprint arXiv:2508.17914},
year = {2025}
}