OutYet reporting
Model ML's finance eval gives GPT-5.6 Sol a file-delivery test
A finance-agent case study evaluates GPT-5.6 Sol on producing reviewable PowerPoint and Excel files, where completion and token use improve in some workflows but correctness and presentation tradeoffs remain.
Model ML, a finance-workflow startup, says it has expanded its use of GPT-5.6 Sol for agents that take a brief and source material through research, analysis, calculations, and into editable PowerPoint decks or Excel workbooks. The company describes a workflow in which a core agent selects tools, reconciles evidence, and routes work to a model, often Sol, while its own document tooling produces native files with traceable sources. That makes this a useful test of the last mile of agent work: not merely whether a model can describe an analysis, but whether a reviewer receives an artifact that can be opened, checked, and edited.
Model ML's PowerPoint results favor Sol on completion and its professional-readiness gate. It reports that Sol produced a .pptx in 100% of its test cases, compared with 76% for Opus 5, and cleared the readiness gate in 43.3% of cases versus 26.7%. Its overall readiness-gated deck-quality score was 59.9%, narrowly above Opus 5 at 56.7% and Fable 5 at 59.3%. The comparison is not a clean sweep: Opus 5 led Sol on the evaluation's visual-quality, layout, and chart-legibility measures, so teams that value presentation polish should not reduce the findings to a single ranking.
Efficiency is also mixed once the work moves from slides to spreadsheets. In Model ML's Excel evaluation, Sol used 2.44 million tokens per workbook against 3.83 million for Opus 5, and averaged seven minutes against 7.5 minutes. But Sol's rate of fully correct models was 50%, below Opus 5 and Fable 5 at 60%, even though its headline key-output accuracy was 83.3%, slightly above Opus 5's 82.8%. For technical buyers, that gap is the important practical limitation: an agent can return a structurally valid workbook and still need substantive checking of its outputs.
Scope matters because an August 6 OpenAI product update changed the ChatGPT version of GPT-5.6 Sol for Plus and Pro users, adding a control for how much thought ChatGPT applies to an answer. OpenAI explicitly says that chat-optimized version is available only in ChatGPT and that the Sol version powering Work and Codex was not changed by that update. Model ML identifies its deployment as an API workflow, so the case study should not be read as evidence that a ChatGPT tuning change caused its reported finance results, or as a blanket result for every Sol surface.
The evidence is strongest as a detailed account of one production harness, not as an independently reproduced cross-vendor benchmark. Model ML's reported setup combines Sol with its own agent planning, data integrations, document-editing tools, code execution environments, and visual review of slides. Its Composite evaluation also tests properties that ordinary text benchmarks often skip, including formulas, source traceability, editability, and file structure. That is why the report is useful for teams building document agents, while also setting a clear limitation: matching these outcomes requires evaluating the model together with the tools, review gates, and artifacts in the intended workflow.