arXiv · 2019 · Preprint

Ardila et al. (Mozilla) · → Paper · Demo: ? · Code: ?

Common Voice is a crowd-sourced, massively-multilingual speech corpus released under a CC0 public domain license, covering 29+ languages with 2,500 hours of transcribed and validated audio as of its November 2019 release.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

Common Voice is a large-scale, openly licensed multilingual speech corpus collected via crowdsourcing, with contributors recording and peer-validating read-speech clips across dozens of languages. The corpus covers a broad typological range, including low-resource languages (Welsh, Breton, Kabyle, Chuvash) that are underrepresented in most other public corpora. TTS and VC researchers use Common Voice as a source of multilingual training and evaluation data, particularly for zero-shot and multilingual synthesis tasks where diverse speaker and language coverage is required. Its permissive CC0 license makes it one of the few large-scale multilingual datasets directly usable for commercial and research model training.

Wiki Connections

multilingual-tts

Cited by: 2210.13438 (EnCodec, uses Common Voice in training data) · 2212.04356 (Whisper, uses Common Voice 5.1 as evaluation benchmark)