What Makes GPT-Realtime-2 Different From Previous Voice Models?
GPT-Realtime-2 is OpenAI’s first voice model with GPT-5-class reasoning capabilities. It can handle multi-step instructions, call tools, handle interruptions, and maintain natural conversational flow without the cognitive lag typical of voice-to-text-to-voice pipelines. A 128K-token context window supports longer, more coherent sessions.
How Does the New Voice Architecture Work?
Instead of routing every interaction through one massive model, developers can now route specific tasks to specialized models:
- GPT-Realtime-2 for conversational reasoning and complex task orchestration
- GPT-Realtime-Translate for live multilingual translation (70+ input languages, 13 output languages)
- GPT-Realtime-Whisper for low-latency streaming transcription
What Are the Pricing and Availability Details?
GPT-Realtime-2 is priced at $32 per 1M audio input tokens and $64 per 1M audio output tokens. GPT-Realtime-Translate costs $0.034 per minute. GPT-Realtime-Whisper costs $0.017 per minute. All models are available in the Realtime API.