arXiv · 2018 · Preprint

Du et al. · → Paper · Demo: ? · Code: ✓

AISHELL-2 releases 1,000 hours of clean Mandarin read speech recorded on iOS, free for academic use, together with a Kaldi-based recipe covering Chinese word segmentation, a two-layer phoneme dictionary (DaCiDian), and a TDNN-based acoustic model achieving 8.81% CER on the iOS test set.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

AISHELL-2 provides a large-scale, publicly available Mandarin speech corpus that enables training of Mandarin ASR components used as intelligibility evaluators in Chinese TTS and multilingual TTS pipelines. The corpus covers 1,991 speakers spanning varied ages, regional accents, and recording environments, offering diversity for training robust speaker-independent acoustic models. The accompanying Kaldi recipe introduces DaCiDian, a two-layer dictionary that decouples word-to-pinyin and pinyin-to-phoneme mappings, making it straightforward for TTS system builders to adapt lexicons to custom phone sets. With iOS, Android, and Microphone channel recordings available for evaluation, the dataset also supports cross-channel acoustic robustness studies relevant to speech synthesis evaluation in diverse deployment conditions.

Wiki Connections

  • multilingual-tts — AISHELL-2 provides the Mandarin training data that multilingual TTS systems rely on for Chinese-language acoustic modelling and intelligibility evaluation.