arXiv · 2025 · Preprint

Comanici et al. (Google) · → Paper · Demo: ? · Code: ?

Introduces the Gemini 2.X model family (2.5 Pro, 2.5 Flash, 2.0 Flash, and 2.0 Flash-Lite), sparse mixture-of-experts transformers with native multimodal support for text, audio, images, and video, including integrated controllable TTS and native audio dialog capabilities.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

Gemini 2.5 is directly relevant to the speech generation community because it integrates audio generation as a first-class output modality. The Gemini 2.5 Preview TTS models support controllable speech synthesis across 80+ languages, with style, emotion, and pacing specified via free-form prompts, and can generate multi-speaker audio for applications such as podcast creation. A separate Native Audio Dialog model enables full-duplex spoken dialogue with style and accent control, turn-taking awareness, and tool use, positioning it as a foundation model for spoken conversational agent systems. TTS research and spoken conversational agent development draw on this model family as a multimodal backbone that provides both language understanding and native audio generation within a single architecture.

Wiki Connections

spoken-language-model · instruction-conditioned-tts · multilingual-tts

In-corpus papers cited by this work: 2407.21783