arXiv · 2021 · Preprint

Chen et al. · → Paper · Demo: ? · Code: ✓

GigaSpeech introduces a 10,000-hour multi-domain English speech corpus drawn from audiobooks, podcasts, and YouTube, along with a scalable forced-alignment and segmentation pipeline for constructing high-quality ASR training data.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

GigaSpeech provides a large-scale, acoustically diverse English speech corpus covering read and spontaneous speaking styles across 24 topic categories. Its five training subsets (10h to 10,000h) and human-verified evaluation sets make it suitable for training and benchmarking ASR systems across a range of data budgets. GigaSpeech serves as a training source for ASR-based intelligibility evaluators (typically Whisper or similar models fine-tuned on it) and as a large-scale pretraining corpus for building speech representations. Its multi-source, multi-style design provides broader acoustic coverage than audiobook-only corpora such as LibriSpeech, which is relevant for TTS systems targeting out-of-domain robustness.

Wiki Connections

evaluation-metrics