English

Human Voice Pitch Estimation: A Convolutional Network with Auto-Labeled and Synthetic Data

Sound 2023-12-19 v2 Machine Learning Audio and Speech Processing

Abstract

In the domain of music and sound processing, pitch extraction plays a pivotal role. Our research presents a specialized convolutional neural network designed for pitch extraction, particularly from the human singing voice in acapella performances. Notably, our approach combines synthetic data with auto-labeled acapella sung audio, creating a robust training environment. Evaluation across datasets comprising synthetic sounds, opera recordings, and time-stretched vowels demonstrates its efficacy. This work paves the way for enhanced pitch extraction in both music and voice settings.

Keywords

Cite

@article{arxiv.2308.07170,
  title  = {Human Voice Pitch Estimation: A Convolutional Network with Auto-Labeled and Synthetic Data},
  author = {Jeremy Cochoy},
  journal= {arXiv preprint arXiv:2308.07170},
  year   = {2023}
}
R2 v1 2026-06-28T11:55:11.330Z