arXiv · 2026 · Preprint

Aniket Deroy (Indian Institute of Technology, Delhi) · → Paper · Demo: ? · Code: ✓

Evaluates Google’s Gemini 2.5 Pro and Flash TTS models on synthesizing persona-driven courtroom advocacy speech across five Indic languages, using a natural-language prompting framework and an 11-dimension human evaluation rubric.

Problem

State-of-the-art TTS systems have reached high naturalness benchmarks for English and other high-resource languages, but generating expressive, professionally authoritative speech for specialized domains such as legal advocacy remains largely unexamined, particularly in linguistically diverse settings such as India. Existing multilingual Indic TTS efforts (BharatGen, Bhashini, IndicTrans2, IndicSynth) have focused on general intelligibility and coverage of the 22 scheduled languages, but none evaluate whether a TTS system can convey the authoritative tone, rhythmic emphasis, and persona-specific “voice” that legal advocacy demands. The paper asks whether a general-purpose multilingual TTS system, guided only by natural-language prompting, can produce courtroom speech that reads as an authentic advocate persona rather than generic narration, across five typologically distinct Indic languages (Hindi, Gujarati, Tamil, Telugu, Bengali).

Method

The study proposes a prompting framework rather than a new model. It defines a persona space in which each of 25 synthetic advocate profiles is characterized by a target language, an area of legal expertise, and a rhetorical style; a large language model generates domain-specific legal argument text conditioned on each profile using a fixed persona instruction template covering identity, linguistic register, courtroom voice, rhetorical strategy, cultural nuance, and discourse structure (exordium, argument, peroration) (§3.1, Table 1). This synthetic legal text is then passed unmodified to Google’s Gemini 2.5 Pro TTS and Gemini 2.5 Flash TTS models, which the paper formalizes as a black-box transformation function mapping text, profile-derived “prosodic steering parameters,” and a language-specific embedding to an output waveform (§3.1.2); no model weights are trained or fine-tuned, and both TTS systems are used purely through prompting. Five advocate profiles are authored per language (25 total), each with a per-language legal-argument prompt (Table 2) and a biography-style persona description (Tables 3-7). Output speech is scored by human evaluators on an 11-dimension, 1-5 Likert Human Evaluation Metric (HEM) covering Naturalness, QMOS (speaking quality), Coherence, Comprehensiveness, Professionalism, Authenticity, Safety, Directiveness, Exploratoriness, Supportiveness, and Expressiveness (§3.1.3, §4).

Key Results

Tables 8 and 9 report per-language HEM scores for the Pro and Flash models respectively. Hindi is the strongest language for both models (QMOS 4.6 Pro / 4.5 Flash; Safety up to 4.9 Pro / 4.4 Flash; Comprehensiveness 4.8 Pro / 4.3 Flash). Bengali and Gujarati trail on Authenticity, Expressiveness, and Exploratoriness for both models: Exploratoriness dips to 3.3 (Gujarati, Pro) and 3.4 (Bengali, Pro), and QMOS drops to 3.3 (Bengali, Flash), the lowest value in either table. Across both model tiers and all five languages, “surface” delivery metrics (Safety 4.2-4.9, Professionalism 4.1-4.8, Comprehensiveness 4.0-4.8) consistently score above “persuasive” metrics (Authenticity 3.4-4.3, Expressiveness 3.4-4.3, Exploratoriness 3.2-4.4), a pattern the paper labels “monotone authority.” No comparison is made against non-Gemini TTS baselines, human advocate recordings, or automated metrics such as WER or PESQ; the evaluation is entirely relative, across languages and between the two Gemini model tiers.

Novelty Assessment

The contribution is a prompting methodology and a domain-specific human evaluation protocol, not a new TTS architecture, training recipe, or reusable dataset. No model is trained or fine-tuned; the paper’s “mathematical framework” formalizes profile-conditioned generation and evaluation as functions but does not correspond to an implemented mechanism beyond prompt construction, so it serves as descriptive notation rather than a novel steering mechanism. The genuinely new element is the domain framing: applying persona-conditioned prompting and an 11-dimension evaluation rubric (spanning both standard naturalness/quality axes and legal-discourse-specific axes such as directiveness and supportiveness) to a previously unstudied use case, courtroom advocacy speech, across five Indic languages. The paper does not report the number or qualifications of human evaluators, inter-rater agreement, or evaluator independence from the author, which limits how much weight the reported scores can carry.

Field Significance

low - the paper documents a narrow, single-author case study of two closed commercial TTS systems on a bespoke, unreleased evaluation set, using an ad hoc rubric without inter-rater reliability reporting. It does not introduce a new architecture, training method, or public dataset, and provides no comparison against non-Gemini baselines. Its main value is as a domain-application data point: current commercial multilingual TTS can deliver clear, “safe” procedural legal speech but shows a measurable gap in expressive, persuasive delivery, with lower scores in Bengali and Gujarati than in Hindi, Tamil, and Telugu.

Claims

  • supports: Prompting-only persona conditioning, natural-language descriptions of professional role, rhetorical style, and register, can differentiate a general-purpose multilingual TTS system’s output style across many distinct personas without any model fine-tuning.

    Evidence: 25 advocate profiles spanning 5 languages and varying rhetorical styles (aggressive/assertive to empathetic/analytical) were each rendered through Gemini 2.5 Pro/Flash TTS using only a fixed instruction template (persona, linguistic identity, voice register, rhetorical strategy), with no model retraining. (§3, Table 1, Table 2)

  • complicates: Human-perceived TTS quality in a multilingual model is not uniform across typologically related target languages, and shared model provenance does not guarantee equal performance.

    Evidence: Within the same evaluation, Hindi scores highest on nearly every HEM dimension (e.g. QMOS 4.6 Pro / 4.5 Flash), while Bengali and Gujarati trail consistently, with QMOS as low as 3.3 for Bengali under Flash, despite all five languages being served by the same underlying Gemini 2.5 models. (§4, Table 8, Table 9)

  • complicates: A single naturalness or quality score can obscure a persistent gap between technical/procedural delivery competence and expressive, persuasive delivery competence in synthesized speech.

    Evidence: Across both Gemini model tiers and all five languages, Safety and Professionalism scores (4.1-4.9) are consistently 0.5-1.5 points higher than Authenticity, Expressiveness, and Exploratoriness scores (3.2-4.3) for the same audio samples. (§4, Table 8, Table 9)

  • complicates: Human evaluation protocols for domain-specific synthetic speech require disclosed evaluator counts, qualifications, and inter-rater reliability to support their conclusions.

    Evidence: The paper defines its Human Evaluation Metric as an average over N raters’ scores but does not report N, rater qualifications, or agreement statistics anywhere in the text. (§3.1.3)

Limitations and Open Questions

The human evaluation methodology is underspecified: the paper defines HEM as an average of N raters' 1-5 scores but never reports how many evaluators participated, their qualifications (e.g. legal or linguistic expertise, native-speaker status per language), or any inter-rater reliability statistic. Given the single-author byline, it is unclear whether the "human evaluation" was independent of the author.

Additional limitations: the study evaluates only two closed, commercial TTS systems (Gemini 2.5 Pro/Flash) with no open-source or non-Gemini baseline for comparison; the legal argument texts are themselves LLM-generated rather than drawn from real courtroom transcripts, so the linguistic and legal authenticity of the input text is not independently verified; no automated intelligibility or speaker-fidelity metrics (WER, PESQ, SPK-SIM) are reported alongside the subjective scores; and the paper’s “mathematical framework” (persona space, TTS transformation function) is descriptive notation rather than an implemented or ablated mechanism, so it provides no evidence about how persona conditioning actually operates inside the black-box Gemini models.

Wiki Connections

  • Multilingual TTS — evaluates a commercial multilingual TTS system’s persona-conditioned output quality across five typologically diverse Indic languages, surfacing a Hindi-favoring performance gap.
  • Instruction-Conditioned TTS — uses natural-language persona instructions (professional role, rhetorical style, courtroom register) as the sole conditioning mechanism for style control, without any model fine-tuning.
  • Evaluation Metrics — introduces an 11-dimension Human Evaluation Metric (HEM) tailored to domain-specific persona synthesis, distinguishing procedural delivery quality from persuasive/expressive quality.
  • Subjective Evaluation — relies entirely on human Likert-scale ratings across 11 dimensions rather than automated metrics, though it does not report evaluator count or inter-rater reliability.