arXiv · 2020 · Preprint

Wang et al. (Facebook AI) · → Paper · Demo: ? · Code: ✓

CoVoST 2 is a large-scale multilingual speech-to-text translation corpus covering translations from 21 languages into English and from English into 15 languages, released under a CC0 license with open-source baselines for ASR, MT, and speech translation.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

CoVoST 2 provides approximately 2,880 hours of speech across 36 translation directions, spanning high-resource European languages (French, German, Spanish) and low-resource languages (Mongolian, Tamil, Estonian, Latvian), sourced from the Common Voice crowdsourced recordings and translated by professional translators. The corpus is built on a Transformer encoder-decoder architecture and includes both cascaded and end-to-end speech translation baselines implemented in fairseq, giving the research community a reproducible benchmark starting point across all supported language pairs. For multilingual speech generation systems that target cross-lingual capabilities, CoVoST 2 provides a training and evaluation resource that spans diverse accent groups (66 in total), age groups, and genders, enabling research into low-resource and massively multilingual settings. The CC0 license places no restrictions on use, making it suitable as a training source for multilingual TTS or speech-to-speech translation pipelines that require paired speech and text across many languages.

Wiki Connections

  • multilingual-tts — CoVoST 2 provides paired speech and text translation data across 21 source and 15 target languages, offering a resource for training multilingual speech generation systems that require cross-lingual supervision.