arXiv · 2024 · Preprint

OpenAI (OpenAI) · → Paper · Demo: ? · Code: ✗

GPT-4o is an end-to-end autoregressive omni model trained jointly across text, audio, image, and video that enables real-time speech-to-speech conversation with human-like response latency (232–320 ms).

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

GPT-4o introduces a natively multimodal autoregressive architecture in which all input and output modalities, including audio, are processed by a single shared neural network rather than by cascaded pipeline components. This end-to-end design allows the model to preserve paralinguistic information across the full conversational turn, enabling real-time spoken dialogue with sub-400 ms latency. TTS and SCA papers cite GPT-4o as a reference point for native speech-in/speech-out capability and as a demonstration that end-to-end training across modalities is feasible at production scale. The system card documents specific safety mitigations for audio generation, including voice output classifiers that enforce compliance with preset voices and post-training procedures to prevent unauthorized voice cloning, which informs later work on controllable and safe TTS deployment.

Wiki Connections

spoken-language-model · speech-to-speech

2502.11946 — Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction