GPT-Realtime-2 is OpenAI's realtime voice model release for low-latency speech-to-speech applications. The first-party model page describes it as supporting configurable reasoning effort, stronger instruction following, reliable tool use, text/audio/image input, and text/audio output, with a 128,000-token context window.
Pricing
OpenAI also publishes separate text-token rates and cached-input discounts for the realtime model on its official pricing page, but I could not verify those values directly in-session.
OpenAI's official pricing/docs for the realtime model indicate token-based multimodal billing. The exact public SKU name 'GPT-Realtime-2' is not separately surfaced in the docs I could access here, so this entry maps it to the current OpenAI realtime pricing. The recorded rates reflect the realtime model's audio-token pricing.