arXiv · 2023 · Preprint

OpenAI (OpenAI) · → Paper · Demo: ? · Code: ✗

GPT-4 is a large-scale multimodal Transformer trained via next-token prediction and RLHF post-alignment, achieving human-level performance on a wide range of professional and academic benchmarks.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

GPT-4 is a decoder-only Transformer pre-trained on text and image inputs with RLHF post-training alignment; its architecture and scale are not disclosed in the technical report. TTS and SCA systems reference GPT-4 as a text understanding and instruction-following backbone, particularly for generating expressive or semantically rich textual representations that guide speech synthesis. Its strong instruction-following capability makes it relevant to instruction-conditioned TTS systems that require a large language model to interpret natural language style prompts before generating speech. Spoken conversational agent research also cites it as a reference point for the quality of language understanding and response generation that audio-native systems aim to match or integrate.

Wiki Connections

spoken-language-model · instruction-conditioned-tts