arXiv · 2025 · Preprint
DeepSeek-AI et al. (DeepSeek) · → Paper · Demo: ? · Code: ✓
DeepSeek-R1 demonstrates that chain-of-thought reasoning capabilities in large language models can be developed through pure reinforcement learning without human-annotated reasoning trajectories, using a multi-stage pipeline combining GRPO-based RL, rejection sampling, and supervised fine-tuning.
Citation Stub
This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.
Context in Speech Generation
DeepSeek-R1 introduces a reasoning-optimised large language model trained via reinforcement learning on a rule-based reward signal, producing extended chain-of-thought responses with self-reflection and verification behaviours. The model is built on the DeepSeek-V3-Base foundation and released in multiple sizes, including smaller distilled variants. Spoken conversational agent and instruction-conditioned TTS systems can draw on DeepSeek-R1 as a language model backbone for tasks requiring multi-step planning, instruction understanding, or dialogue reasoning. The RL training methodology and the GRPO algorithm it employs are also referenced by speech generation papers exploring reinforcement learning for alignment and quality optimisation of TTS outputs.
Wiki Connections
Related concepts: rlhf-speech · spoken-language-model
In-corpus papers cited by this work: 2412.19437 · 2402.03300 · 2307.09288