arXiv · 2020 · Preprint

Pratap et al. (Facebook AI Research) · → Paper · Demo: ? · Code: ?

Introduces Multilingual LibriSpeech (MLS), a freely available multilingual speech corpus derived from LibriVox audiobooks covering 8 languages with 44.5K hours of English and approximately 6K additional hours across German, Dutch, French, Spanish, Italian, Portuguese, and Polish.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

MLS provides one of the largest openly available multilingual read-speech corpora, constructed from LibriVox audiobooks via an automated pipeline combining ASR-based segmentation, TF-IDF transcript retrieval, and Smith-Waterman alignment. The dataset is split to avoid speaker overlap between train, dev, and test partitions, and includes human-verified dev and test transcripts across all eight languages. TTS research uses MLS as a multilingual training resource for building and evaluating speech synthesis systems that require broad language coverage, particularly for multilingual or cross-lingual zero-shot and speaker-conditioned generation. The corpus is freely available at openslr.org.

Wiki Connections

multilingual-tts

Cites in-corpus: 1904.02882

2402.01912 — Natural language guidance of high-fidelity text-to-speech with synthetic annotations