OutYet reporting
GPT-5.6 Sol’s latest change is a ChatGPT product split, not a new developer model
OpenAI’s August update changes the ChatGPT behavior and controls around GPT-5.6 Sol, while leaving its Work and Codex version untouched. A new finance-workflow case study offers useful but vendor-mediated evidence about where the model can help and where teams should keep their own evaluation gates.
OpenAI’s August 6 update changes the ChatGPT-specific version of GPT-5.6 Sol for Plus and Pro users: the company says it has tuned the model toward more focused answers and improved factual reliability, and added a slider that lets users select how much thought a response receives. The update is explicitly scoped to ChatGPT. OpenAI says the GPT-5.6 Sol version used in Work and Codex is not changing, so developers should not treat the announcement as evidence of an API, Codex, or enterprise-runtime revision.
The comparison OpenAI presents is primarily against GPT-5.5 Instant in conversational use. Its examples emphasize answering the central question before supplying detail, and OpenAI reports an internal evaluation in which responses with at least one factual error were less common for GPT-5.6 Sol than for GPT-5.5 Instant on financial, medical, and legal prompts. That is a product claim and an internal measurement, not an independent benchmark; it is useful context for the intended behavior change, but it does not establish performance across external workloads.
A separate OpenAI customer story gives a more concrete, though still vendor-mediated, view of GPT-5.6 Sol inside an agent harness. Model ML says it evaluated the model on finance assignments that end in editable PowerPoint decks and Excel workbooks, with checks for numbers, sources, formulas, structure, and presentation quality. Its published table reports stronger deck completion and professional-readiness rates for Sol than for Opus 5 in that setup, while also showing mixed tradeoffs: Sol trailed Opus 5 on some visual-quality measures and did not lead on fully correct Excel models.
For technical teams, the practical takeaway is to separate the two claims. The ChatGPT update may change interactive behavior and reasoning control for Plus and Pro users, but OpenAI says it does not alter the Work and Codex version. Model ML’s results suggest that Sol can be worth testing where an agent must produce editable, reviewable office files, yet the evidence comes from one company’s benchmark, toolchain, and scoring rules. Teams should reproduce representative tasks, validate source traceability and numerical outputs, and measure token and latency costs in their own harness before changing routing decisions.