arXiv · 2025 · Preprint
An Yang et al. · → Paper · Demo: ? · Code: ✓
Qwen3 introduces a family of open-weight LLMs spanning 0.6B to 235B parameters (dense and MoE) with a unified thinking/non-thinking mode framework, 36 trillion token pre-training, and support for 119 languages.
Citation Stub
This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.
Context in Speech Generation
Qwen3 provides the language model backbone used by TTS and SCA systems that require a capable, open-weight text encoder or autoregressive generator. The model family’s multilingual coverage across 119 languages and dialects makes it particularly useful for multilingual TTS pipelines where the LM must handle diverse scripts and linguistic patterns. Its range of model sizes (0.6B to 235B) enables researchers to select a Qwen3 variant suited to their compute budget, and the Apache 2.0 license permits unrestricted use in research and commercial speech systems. The integration of thinking and non-thinking modes in a single model also makes Qwen3 applicable to instruction-conditioned and agent-driven speech generation settings where the LM component must handle both rapid responses and complex multi-step reasoning.
Wiki Connections
Related concepts: multilingual-tts, spoken-language-model
In-corpus papers cited by Qwen3: 2309.16609, 2407.21783, 2501.12948, 2412.19437, 2402.03300, 2503.19786, 2407.10671, 2412.15115