OutYet reporting

GPT-5.6 Sol gets a ChatGPT tune-up, while a finance case study shows the limits of headline benchmarks

OpenAI’s latest Sol update is a ChatGPT product change, while a finance-agent case study offers useful but vendor-hosted evidence about document workflows.

OutYet Editorial Desk

OpenAI's August 6 product update changes GPT-5.6 Sol inside ChatGPT for Plus and Pro users. The company says it retuned the model for more focused answers and better factual reliability, and added a slider that lets users choose how much thought a response receives. OpenAI also says the update is confined to ChatGPT: the GPT-5.6 Sol versions used in Work and Codex are not changing.

That boundary is important context for a model OpenAI previewed on June 26. The August notice describes a ChatGPT update and distinguishes it from a separate access change: GPT-5.6 Luna is to become the default for Free and Go users, with unlimited text chats and a Think button planned for the following week. The notice explicitly says the Sol versions in Work and Codex are unchanged, so this evidence should not be read as a change to those product surfaces.

A second August item provides a concrete, narrower look at how GPT-5.6 Sol performs in an agentic document workflow. OpenAI says Model ML routes finance tasks through research, calculations, and native PowerPoint or Excel generation, with a core agent selecting tools and models. Model ML's Composite evaluation, as reported by OpenAI, measures output correctness, source and formula checks, structure, and presentation quality rather than treating a text-only answer as the finished deliverable.

The reported comparison is mixed in ways that matter more than a single ranking. In Model ML's PowerPoint tests, Sol produced a .pptx in 100% of cases versus 76% for Opus 5 and cleared the firm's professional-readiness gate in 43.3% of cases versus 26.7%. Yet the same table gives Sol lower visual-quality and chart-legibility scores than Opus 5. In Excel, Sol used fewer tokens per workbook than Opus 5, but its fully-correct-model rate was 50.0% versus 60.0% for both Opus 5 and Fable 5.

For technical users, the practical signal is specific: ChatGPT users gain a control over how much effort Sol uses, while teams building document agents get a vendor-hosted example of performance inside a substantial harness. Model ML attributes its workflow to document tools, data integrations, code execution, toolkit loading, and human review of assumptions and sources. That makes the reported results useful evidence for evaluating an end-to-end system, but not a substitute for testing a team's own templates, tools, review gates, and workloads.

Related models

Sources