OutYet reporting
GPT-5.6 Sol gets an API preview built around latency, not a new model
OpenAI's Ultrafast tier pairs GPT-5.6 Sol with Cerebras infrastructure, but limited availability and provider-reported performance keep it a deployment preview rather than a broad model change.
OpenAI is previewing Ultrafast, an API service tier for GPT-5.6 Sol that it says can run up to 14 times faster than Standard processing and produce up to 750 output tokens per second. The company describes it as a limited preview for select customers, with expansion tied to available capacity. That makes this a serving change around an existing model, not evidence of a separate model launch or a general-access commitment. OpenAI says Cerebras powers the tier.
The timing distinguishes this announcement from OpenAI's earlier ChatGPT update. On August 6, OpenAI updated GPT-5.6 Sol in ChatGPT for Plus and Pro users, emphasizing focused answers, fact reliability, and a response-effort slider. It also said that version was limited to the Chat experience and that the GPT-5.6 Sol version powering Work and Codex was unchanged. Ultrafast begins with API access, so teams should not infer that the ChatGPT update has changed their API or Codex deployment.
The comparison with GPT-5.5 is not only about faster generation. OpenAI's builder guide says GPT-5.6 Sol at low reasoning outperformed GPT-5.5 at high reasoning on its Agents' Last Exam evaluation when the harness was held constant. The same guide argues that retained reasoning, native compaction, multi-agent orchestration, and programmatic tool calling can reduce avoidable context work. Ultrafast adds a different lever: faster serving for the same flagship model, rather than a claim that latency alone improves answer quality.
For technical users, the immediate use case is a latency-sensitive workflow in which a useful next step loses value if it arrives too late, such as incident triage or an interactive research loop. OpenAI's examples still retain an important limitation: engineers remain responsible for judgment and deployment. The speed figures are provider-reported, the preview is capacity-limited, and OpenAI has not represented this announcement as a change to the Work or Codex model. Evaluation should therefore happen in an existing workload harness, with attention to end-to-end latency and output quality rather than output tokens per second alone.