arXiv · 2025 · Preprint

Mandai et al. (Institute of Science Tokyo / AIST / University of Amsterdam) · → Paper · Demo: ? · Code: ?

A four-phase user study (N=512) demonstrates that fundamental frequency (F0) and first formant manipulation can amplify perceived kawaiiness in TTS voices, while professionally recorded game character voices resist the same technique due to voice-specific ceiling effects. Published at CHI ‘25.

Problem

Prior research on kawaii (the Japanese concept of cuteness with socioemotional and cross-cultural dimensions) has been almost entirely visual. The small body of work on kawaii vocalics had identified associations between social identity perceptions (girlishness, gender ambiguity, youth) and voice kawaiiness in descriptive studies, but no one had attempted to actively manipulate those perceptions through acoustic signal processing. It was unknown whether fundamental and formant frequency shifts could reliably induce kawaii responses in synthetic voices, and whether any such method would transfer across voice types.

Method

The study proceeds in four phases using existing acoustic manipulation tools, not a novel synthesis system. In Phase 1, five Japanese TTS voices from CoeFont and MAKOTO are manipulated by shifting F0 and formants in increments of three semitones (up or down) using Steinberg Cubase. A parallel study in the same phase compares this manual approach to automated vocoder-based methods: Legacy-STRAIGHT and WORLD, both of which decompose speech into F0, spectral envelope, and aperiodicity for independent manipulation. Participants (n=50 per study) rate each clip on a 7-point Likert kawaii scale along with gender, age, and humanlikeness items.

Phase 2 applies WORLD-based automated shifting to the 18 professionally recorded game character voices from an earlier corpus study, evaluated by n=150 participants. Phase 3 returns to manual Cubase manipulation for those same game voices, this time testing one-, two-, and three-semitone increments to probe for finer-grained sweet spots, evaluated by n=51 participants. Statistical analyses use Spearman correlations, Mann-Whitney U, Friedman, and Wilcoxon signed-rank tests throughout, with Generalized Estimating Equations (GEE) for covariate analysis in Phase 3.

Key Results

For TTS voices (Phase 1), kawaii perceptions correlate strongly and positively with F0 (rs=0.89, p<0.001) and with the first formant F1 (rs=0.74, p<0.001). F2 and F3 do not reach significance for kawaii, though F3 shows a modest positive correlation with gender ambiguity perceptions. Age perceptions likewise decrease with higher F0 and F1, consistent with the hypothesis that younger-sounding voices are perceived as more kawaii.

Comparing automated vocoders to manual editing (Phase 1, study 2), no significant differences in kawaii ratings emerge across Cubase, Legacy-STRAIGHT, and WORLD for any of the five TTS voices (Table 3). However, significant differences appear on secondary attributes: Legacy-STRAIGHT increases perceived humanlikeness for one voice, and both automated methods alter animal-likeness and trustworthiness ratings for certain voices (Table 4).

For game character voices (Phase 2), the automated three-semitone manipulation actually decreases perceived kawaiiness relative to the originals (M=3.30 vs. M=3.44, Mann-Whitney U, p<0.001), rejecting H3. The effect size is small (r=0.07). GEE analysis finds that favorability, humanlikeness, familiarity, and trustworthiness jointly predict kawaiiness — not frequency directly.

Returning to manual manipulation in Phase 3, kawaiiness increases significantly for several specific characters (Toad, Ayaka, Peach) at one or two semitones, while others (Edea, Pikachu) show a ceiling or reversal at three semitones. The scatter of F1 against kawaiiness in Phase 3 reveals a quadratic shape with a vertex, suggesting a sweet spot rather than monotonic improvement.

Novelty Assessment

The paper’s contribution is empirical rather than methodological. WORLD and Legacy-STRAIGHT are established vocoders with decades of use in speech research; shifting F0 and formants by a semitone ratio is a standard operation in both tools. The novelty lies in applying this machinery to systematically measure and manipulate kawaii perceptions across a sizable participant pool and two distinct voice types. The four-phase design provides replication and boundary-condition analysis that the earlier descriptive kawaii vocalics work lacked. The findings of voice-specific sweet spots and ceiling effects are genuinely informative for voice UX design, though the sample is limited to Japanese participants and Japanese-language stimuli.

Field Significance

Low — this paper advances the nascent field of kawaii vocalics rather than core speech synthesis research. It provides a usable, if limited, method for modifying cuteness perceptions of TTS voices through F0 and formant shifts, and documents where that method fails (professionally processed character voices, older adult voices). The findings are most relevant to voice UX practitioners designing synthetic agents intended to project a girlish or youthful character, and to researchers studying the perceptual attributes of voice beyond intelligibility and naturalness.

Claims

  • supports: Fundamental frequency and lower formant frequencies are primary acoustic predictors of perceived cuteness in synthetic voices.

    Evidence: Phase 1 Spearman correlations show strong positive relationships between kawaii ratings and F0 (rs=0.89) and F1 (rs=0.74), while F2 (p=0.08) and F3 (p=0.25) are non-significant. (§4.6.1)

  • complicates: Acoustic frequency manipulations that amplify cuteness perceptions in generative TTS voices do not transfer to naturally recorded or professionally processed voices.

    Evidence: The three-semitone F0/formant shift that improved kawaii for TTS voices (Phase 1) significantly reduced kawaiiness for game character voices in Phase 2 (H3 rejected, U-test p<0.001, r=0.07), attributed to processing artefacts and possible ceiling effects in professionally crafted voices. (§5.4, Table 6)

  • complicates: Automated vocoder-based pitch and formant shifting matches manual audio editing for cuteness perception but diverges on secondary perceptual attributes.

    Evidence: Wilcoxon signed-rank tests found no significant differences in kawaii ratings across Cubase, Legacy-STRAIGHT, and WORLD for all five TTS voices, but significant differences emerged for humanlikeness, animal-likeness, trustworthiness, and excitedness in several voices. (§4.6.2, Table 3, Table 4)

  • supports: Perceived cuteness of synthetic voices is more strongly predicted by social desirability attributes than by acoustic frequency measures in isolation.

    Evidence: GEE analysis across game character voices found favorability (coef=0.51), humanlikeness (coef=0.19), familiarity (coef=0.14), and trustworthiness (coef=0.13) as significant predictors of kawaiiness, with effect sizes that exceed those of frequency correlations. (§5.4, Table 7)

Limitations and Open Questions

Warning

All four study phases recruited exclusively Japanese adults via Yahoo! Crowdsourcing Japan, with no participants under 18 and a participant pool skewed toward ages 35-54. Generalisability of the kawaii vocalics manipulation to non-Japanese listeners, younger audiences, or voices in languages other than Japanese is untested.

The paper lacks a validated measurement scale for kawaii, relying on a single Likert item from prior work. Perceptual ratings across a large number of voice clips per session introduce potential fatigue and order effects that could not be fully controlled despite randomisation. The study treats F0 and formants as independent variables but does not control for interactions between them or for speech content effects. The automated manipulation pipeline (WORLD, Legacy-STRAIGHT) introduces resynthesis artefacts that may confound perceptual ratings, as the authors acknowledge.

Wiki Connections

  • Subjective Evaluation — this paper contributes a four-phase listener study methodology for measuring voice perceptual attributes beyond intelligibility, specifically quantifying kawaii responses to acoustic manipulations.
  • Evaluation Metrics — the paper develops and applies a kawaii vocalics measurement framework using Spearman correlations and GEE analysis to link acoustic features to perceptual dimensions of voice cuteness.
  • Prosody Control — the manipulation technique (semitone-based F0 and formant shifting via WORLD/STRAIGHT) represents an explicit, parameterised mechanism for controlling pitch and vocal tract resonances independently of speech content.