arXiv · 2019 · Preprint

Takamichi et al. · → Paper · Demo: ? · Code: ?

Introduces the JVS (Japanese Versatile Speech) corpus: 30 hours of studio-quality Japanese speech from 100 professional speakers, covering normal, whisper, and falsetto vocal styles, with parallel utterances and rich speaker-level annotations.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

The JVS corpus provides a multi-speaker Japanese training and evaluation resource for TTS and voice conversion research, offering 100 studio-recorded professional speakers at 24 kHz with phoneme alignments, F0 range annotations, and perceptual speaker similarity scores. The 22-hour parallel sub-corpus (100 utterances common across all speakers) is specifically suited for voice conversion experiments, where parallel data enables direct spectral mapping between speakers. The three vocal style sub-corpora (normal, whisper, falsetto) enable speech synthesis research beyond standard reading-style data, including style conversion and multi-style modeling tasks. Speaker similarity matrices derived from crowdsourced perceptual judgments across all same-gender speaker pairs provide a structured resource for speaker space modeling in multi-speaker TTS systems.

Wiki Connections

  • multilingual-tts — the JVS corpus extends Japanese-language TTS resources, complementing Japanese-only corpora used to train and evaluate multilingual or language-specific speech synthesis systems.
  • voice-conversion — the parallel100 sub-corpus provides 100 phonetically balanced utterances shared across all 100 speakers, enabling parallel-data voice conversion experiments.
  • 1609.03499 (WaveNet) — cited in the introduction as one of the deep learning developments that made data-driven speech synthesis feasible, motivating the need for freely accessible multi-speaker corpora.