English

Unsupervised Classification of Voiced Speech and Pitch Tracking Using Forward-Backward Kalman Filtering

Sound 2021-03-02 v1 Machine Learning Audio and Speech Processing

Abstract

The detection of voiced speech, the estimation of the fundamental frequency, and the tracking of pitch values over time are crucial subtasks for a variety of speech processing techniques. Many different algorithms have been developed for each of the three subtasks. We present a new algorithm that integrates the three subtasks into a single procedure. The algorithm can be applied to pre-recorded speech utterances in the presence of considerable amounts of background noise. We combine a collection of standard metrics, such as the zero-crossing rate, for example, to formulate an unsupervised voicing classifier. The estimation of pitch values is accomplished with a hybrid autocorrelation-based technique. We propose a forward-backward Kalman filter to smooth the estimated pitch contour. In experiments, we are able to show that the proposed method compares favorably with current, state-of-the-art pitch detection algorithms.

Keywords

Cite

@article{arxiv.2103.01173,
  title  = {Unsupervised Classification of Voiced Speech and Pitch Tracking Using Forward-Backward Kalman Filtering},
  author = {Benedikt Boenninghoff and Robert M. Nickel and Steffen Zeiler and Dorothea Kolossa},
  journal= {arXiv preprint arXiv:2103.01173},
  year   = {2021}
}

Comments

Speech Communication; 12. ITG Symposium, 5-7 Oct. 2016