Interspeech · 2025 · Conference

Kondo et al. (NTT Corporation) · → Paper · Demo: ? · Code: ?

JIS is a Japanese multi-speaker corpus of 169 live idol voices with publicly identified speakers, designed to enable more rigorous speaker similarity evaluations in TTS and VC by allowing experiment planners to recruit listeners who are genuinely familiar with the speakers.

Problem

Existing multi-speaker corpora for speech generation — LibriSpeech, VCTK, and the Japanese JVS corpus — have two structural limitations for speaker similarity evaluation. First, speakers span a wide range of attribute categories, so listeners can accept approximate attribute matches as adequate speaker similarity even when subtle voice characteristics are not reproduced. Second, all speakers are anonymized, preventing recruitment of listeners who are familiar with them; evaluations therefore rely on naive listeners who cannot detect fine-grained voice differences. Together, these factors risk systematically tolerant subjective evaluations that do not reflect whether generated voices would satisfy users who want to hear a specific, known person’s voice.

Method

JIS contains 169 speakers drawn from 33 Japanese live idol groups. Crucially, each speaker is assigned her commonly used stage name rather than an anonymous identifier, making it possible to recruit fans who know the speaker’s voice. The corpus divides into Speech A (61 speakers, studio-recorded under professional supervision) and Speech B (108 speakers, recorded in quiet but unspecified rooms to expand speaker count within a constrained budget). Both subsets include phoneme-balanced read speech from the voiceactress100 text corpus, spontaneous speech covering idol-activity scenarios (post-performance farewell greetings, photo session phrases, and personal introductions), nine Japanese greeting phrases, and singing (Katatsumuri, a royalty-free folk song also used in JVS-MuSiC). All audio is saved at 48 kHz / 24-bit. Supplementary metadata includes group name, stage name, prefecture of origin, and mutual voice-impression questionnaire responses collected within each idol group.

Audio quality is characterised using UTMOS (mean pseudo-MOS: Speech A 3.4, Speech B 2.8, JVS parallel100 3.7 for reference). Speaker distinctiveness is verified through x-vector and ECAPA-TDNN embeddings, which show that JIS speakers cluster separately from JVS female speakers in t-SNE projections, and that Speech A and B are interleaved within the JIS cluster, indicating that recording environment differences do not strongly affect speaker identity representations.

Key Results

UTMOS-based quality assessment places Speech A below JVS parallel100 (3.4 vs. 3.7), attributed to linguistic mismatch with UTMOS training data and hesitations from non-professional speakers rather than recording conditions. Speech B shows lower average quality (2.8) and higher variance consistent with its more varied recording environments. Speaker embedding analysis demonstrates that speaking style systematically shifts embedding position within a speaker cluster: post-performance greeting speech and photo-session speech produce concentrated style-driven clusters across speakers, while self-introduction speech shows more dispersed, individual-specific patterns.

Novelty Assessment

The contribution is a corpus, not a model or training method. The distinguishing feature relative to existing Japanese corpora is the non-anonymous design: using stage names for publicly active entertainers makes familiarity-based listening experiments feasible for the first time in Japanese speech generation research. The analysis sections confirm speaker distinctiveness and surface useful quality caveats, but introduce no methodological novelty. The corpus is small by current standards (169 speakers, 17 hours), restricted to non-commercial use, and limited to young Japanese female voices, so its applicability as general-purpose training data is narrow.

Field Significance

Moderate — JIS fills a specific methodological gap in speaker similarity evaluation by providing the first non-anonymous Japanese multi-speaker corpus for speech generation research. Its practical value is in enabling rigorous listening experiments with familiar listeners; as a training resource it is limited by size, restricted licensing, and a homogeneous speaker demographic.

Claims

  • supports: Non-anonymous speaker corpora with publicly identifiable voices enable more rigorous subjective evaluations of speaker similarity in TTS and VC systems.

    Evidence: JIS assigns stage names to 169 Japanese live idol speakers, allowing experiment designers to recruit listeners familiar with the speakers, enabling discrimination of subtle voice characteristics that anonymous corpus evaluations cannot capture. (§1, §3.1)

  • complicates: Automatic MOS predictors trained on TTS-generated speech may underestimate audio quality when applied to recordings of non-professional speakers, even under studio conditions.

    Evidence: JIS Speech A (studio-recorded) achieves a mean UTMOS of 3.4 compared to JVS parallel100’s 3.7, with the gap attributed to linguistic mismatch in UTMOS training data and speech hesitations inherent to non-professional speakers rather than recording quality differences. (§4.2.1)

  • supports: Speaking style and communicative context introduce systematic variation in speaker embeddings that interacts with speaker identity, presenting a challenge for style-robust speaker representation.

    Evidence: ECAPA-TDNN embeddings of JIS speakers show that specific speaking styles (energetic post-performance greetings, intimate photo-session speech) produce cross-speaker clusters in t-SNE, partially overriding individual speaker identity, while speech expressing personal individuality is more dispersed. (§4.2.2)

Limitations and Open Questions

The corpus is small (17 hours, 169 speakers) and covers only young Japanese female voices, restricting direct use for general-purpose or multilingual TTS training. Distribution requires a signed agreement and is limited to non-commercial basic research. Speech B recording conditions are unspecified and variable across groups, introducing inconsistencies in audio quality. No TTS or VC model is trained on JIS in this work, leaving empirical validation of the corpus’s utility for model development open.

Wiki Connections

  • Evaluation Metrics — JIS employs UTMOS for audio quality characterisation and demonstrates that TTS-trained MOS predictors can underestimate quality for non-professional speaker recordings, a relevant caveat for automatic evaluation in this domain.
  • Subjective Evaluation — JIS is designed specifically to improve the rigour of subjective speaker similarity listening experiments by replacing anonymous speakers with publicly identifiable stage names, enabling familiarity-based listener recruitment.
  • Zero-Shot TTS — JIS targets evaluation settings where generated voices must match specific known speakers, a core requirement of zero-shot synthesis evaluation.
  • Voice Conversion — the corpus is explicitly intended for VC research alongside TTS, with its non-anonymous design particularly relevant for evaluating whether converted voices satisfy listeners familiar with the target speaker.
  • Speaker Adaptation — the supplementary questionnaire data and multi-style recordings provide resources for studying voice characterisation and preference-based speaker modelling.