English

MUSAN: A Music, Speech, and Noise Corpus

Sound 2015-10-30 v1

Abstract

This report introduces a new corpus of music, speech, and noise. This dataset is suitable for training models for voice activity detection (VAD) and music/speech discrimination. Our corpus is released under a flexible Creative Commons license. The dataset consists of music from several genres, speech from twelve languages, and a wide assortment of technical and non-technical noises. We demonstrate use of this corpus for music/speech discrimination on Broadcast news and VAD for speaker identification.

Keywords

Cite

@article{arxiv.1510.08484,
  title  = {MUSAN: A Music, Speech, and Noise Corpus},
  author = {David Snyder and Guoguo Chen and Daniel Povey},
  journal= {arXiv preprint arXiv:1510.08484},
  year   = {2015}
}