GPT-Realtime-2 brings GPT-5 intelligence to voice API
OpenAI released a new generation of voice models in its API on Wednesday, giving developers tools to build apps that can reason through spoken requests, translate across +70 languages, and transcribe speech as it happens. The three models are named GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. They move AI voice interfaces beyond simple Q&A exchanges into a territory where an AI agent can listen, think, and act mid-conversation. GPT-Realtime-2 is the flagship. OpenAI says it offers GPT-5-class reasoning, a significant step up from its predecessor, GPT-Realtime-1.5. The model scored 15.2% higher on Big Bench Audio, a benchmark for audio intelligence, and 13.8% higher on Audio MultiChallenge, which tests instruction following in multi-turn spoken dialogue. The practical upgrades target developers building production voice agents. The model now supports a 128K context window, quadrupled from the previous 32K limit, and offers five tiers of adjustable reasoning effort from “minimal” to “xhigh.” It can call multiple tools simultaneously, recover from errors with spoken acknowledgments, and produce short bridging phrases like “let me check that” while processing a request. GPT-Realtime-Translate handles live speech translation. It accepts more than 70 input languages and outputs in 13, designed to keep pace with a speaker in real time. GPT-Realtime-Whisper provides streaming speech-to-text (STT), transcribing words as