arXiv · 2015 · Preprint

Snyder et al. · → Paper · Demo: ? · Code: ?

Introduces MUSAN, a freely redistributable corpus of approximately 109 hours of music, speech, and noise collected from Creative Commons and US Public Domain sources, designed to train voice activity detection and music/speech discrimination models.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

MUSAN provides a training and augmentation resource comprising roughly 60 hours of multilingual speech (12 languages), 42 hours of music across multiple genres, and 6 hours of assorted noise recordings, all formatted as 16 kHz WAV files. Its Creative Commons and Public Domain licensing makes it suitable for both research and commercial use, distinguishing it from earlier corpora such as GTZAN where redistribution rights were unclear. TTS and speaker modelling pipelines can use the noise and music partitions for data augmentation, simulating adverse acoustic conditions during training. The VAD application demonstrated in the paper, using GMM-based classifiers trained on MUSAN to improve speaker recognition, illustrates its role as an auxiliary training resource rather than a primary speech synthesis dataset.

Wiki Connections

  • evaluation-metrics — MUSAN enables training of VAD and music/speech discrimination components that feed into downstream speech quality and speaker verification evaluations.