OpenAI is previewing a new Ultrafast service tier designed to run GPT-5.6 Sol faster than Standard for API workloads that need quick responses.

What Ultrafast is

OpenAI says Ultrafast can reach up to 14 times the speed of Standard, with output of up to 750 tokens per second. It is a service tier and inference configuration, not a renamed GPT-5.6 Sol model.

Ultrafast is currently a limited preview for selected customers, and expansion depends on available capacity. The published ceiling should not be read as a fixed speed for every account or task.

Where faster output matters

Customer support, voice interaction, real-time research, and agents cannot always make users wait. Faster output can shorten conversational pauses and let a workflow complete more steps in the same time.

OpenAI also points to incident response, financial research, security analysis, and real-time commerce. These are suggested directions to test, not independent proof of performance in every scenario.

Speed is not the only metric

Teams still need to measure cost, rate limits, concurrency, output stability, and model quality. Prompt length, output content, and API conditions can change the observed token rate.

Ultrafast shows how providers are turning different speed tiers for the same model into a product choice. Teams may need to weigh latency alongside quality, price, and context length.

The accurate description is that Ultrafast has been announced and entered a limited preview. It is not generally available, and every GPT-5.6 Sol user will not immediately get 14x speed.

Treat it as a performance option

Ultrafast is worth testing for products that need immediate interaction, while Standard may remain better for batch summaries or background agents. Adoption should follow measurements of latency, cost, and quality on real tasks.