GPT-5.6 Sol cuts token use, lifts pass rates in Model ML's finance decks

GPT-5.6 Sol cuts token use, lifts pass rates in Model ML's finance decks

OpenAI published a customer story about Model ML, a finance-workflow software company founded by brothers Arnie and Chaz Englander. The brothers built the software for themselves after two successful exits, once they began investing through a private family office and needed tools to handle the work. Grown out of that internal software, Model ML's agents now carry a finance professional's request from the initial brief through research and analysis to a finished, editable PowerPoint deck or Excel workbook.

At the center of the system, a core agent plans the work, selects the right tools, reconciles evidence, runs calculations, and routes each step to the model best suited to it, most often GPT-5.6 Sol. One of the tools it calls is Model ML's own document tooling, which produces native PowerPoint and Excel files with sources traced back to the underlying evidence. Model ML calls the product "surface-agnostic": a finance professional can start an assignment by email or in the Model ML app and continue it inside Microsoft Office plug-ins without re-explaining the task. For an investment-committee deck, the agent can turn a brief and source material into an editable PowerPoint; for an Excel task, it can start from a client template or a blank workbook, gather the required data, build formulas and logic across multiple tabs, and apply finance-specific formatting.

On Model ML's Composite benchmark, its own evaluation suite for AI in financial services that follows an assignment from brief through research and calculations to a finished deck or spreadsheet and then checks the numbers, sources, formulas, structure and visual quality, GPT-5.6 Sol used about 21% fewer tokens per PowerPoint deck than Fable 5 and 36% fewer tokens per Excel workbook than Opus 5. For PowerPoint specifically, an evaluation spanning hundreds of generated decks, GPT-5.6 Sol completed the workflow in 100% of test cases against 76% for Opus 5, and cleared Model ML's professional-readiness gate, a measure of whether the output was ready for substantive review, in 43.3% of cases versus 26.7% for Opus 5, a 16.6-percentage-point lead. It also led Opus 5 on deck quality, brief adherence, hierarchy and consistency.

Those results led Model ML to expand GPT-5.6 Sol into production, including some workflows that had previously run on Opus 4.8. The company points to effects in daily work: at one global asset manager, a bespoke tearsheet that used to take an analyst about an hour to assemble now takes about five minutes. In a separate workflow, Model ML's agents processed virtual data rooms containing more than 100,000 rows and hundreds of files in a single pass.

Model ML runs GPT-5.6 Sol inside its own agent harness, paired with document-creation and editing tools that let the agent build slides with editable graphs and tables. The harness keeps the original brief in context throughout the task and reviews every slide visually before returning the finished file. Model ML reached this setup through on-site sessions with OpenAI, in which the two teams traced how the agent planned presentations, selected tools and maintained context, then used those findings to refine the agent's instructions and decide when each toolkit should load.

Model ML's customers are moving toward browser-based outputs that stay connected to the underlying models and source material: the platform can generate secure, interactive outputs that update continuously or lock to a moment in time, letting a reviewer open an investment summary, click through to the financial model behind a figure, and keep working with the agent on the same page. As Englander put it: "Ready for real work means the user can move directly into real review. The numbers trace back, the workbook recalculates, the slide is editable." Two other lines in the piece are not attributed to a named speaker: one contrasts GPT-5.6 Sol with earlier models that needed the task broken down in detail before producing a close-to-final output, the other frames the user's role as focusing on judgment, such as refining assumptions or sharpening the message, rather than rebuilding the analysis.

Key facts

  • GPT-5.6 Sol used about 21% fewer tokens per PowerPoint deck than Fable 5, and 36% fewer tokens per Excel workbook than Opus 5, on Model ML's Composite benchmark.
  • GPT-5.6 Sol completed the PowerPoint workflow in 100% of test cases versus 76% for Opus 5, and cleared Model ML's professional-readiness gate in 43.3% of cases versus 26.7% for Opus 5, a 16.6-percentage-point lead.
  • At one global asset manager, a bespoke tearsheet that took an analyst about an hour to build now takes about five minutes using Model ML's agents.
  • In another workflow, Model ML's agents processed virtual data rooms containing more than 100,000 rows and hundreds of files in a single pass.
  • The Composite results led Model ML to expand GPT-5.6 Sol into production, including some workflows previously handled by Opus 4.8.

Why it matters

Finance work has a demanding last mile before it can go in front of clients or senior decision-makers: reconciling evidence, building and formatting the file, checking every number, and linking each claim back to its source. This case study is OpenAI's evidence that GPT-5.6 Sol handles more of that last mile unsupervised than the models Model ML compared it against, completing the PowerPoint workflow outright far more often and using notably fewer tokens per file, which OpenAI frames as agents needing less hand-holding to reach a usable final output.

Who it affects

Finance professionals who build decks and workbooks for clients, investment committees or deal teams are the direct users described. Model ML is the vendor whose product and benchmark generated the numbers, and OpenAI is positioning GPT-5.6 Sol against rival models, named here as Opus 5, Opus 4.8 and Fable 5, for agentic, tool-using finance work rather than plain chat.

How to use it

The workflow described runs through Model ML's own product, not a generic GPT-5.6 Sol integration: a finance professional starts an assignment by email or in the Model ML app and continues it inside Microsoft Office plug-ins, with Model ML's agent harness handling tool selection, document creation and slide review. The source gives no pricing, availability date or rollout timeline for GPT-5.6 Sol itself.

How solid is it

This is OpenAI's own customer story, and every comparative figure comes from Model ML's Composite benchmark, a suite the company built and runs itself rather than an independent or third-party evaluation. The source gives no sample size or methodology beyond noting the PowerPoint evaluation spans hundreds of generated decks, and does not say whether any of the results were checked outside Model ML's own scoring.

Risks and caveats

The comparison is self-reported by the vendor and the model maker together, with no outside verification described. Even on GPT-5.6 Sol, the professional-readiness gate passed in only 43.3% of cases, meaning most generated decks still needed further work before they were ready for substantive review. Two of the four quotes in the piece are not attributed to a named speaker, and the source does not specify which of the two Englander brothers is behind the other two.

“Ready for real work means the user can move directly into real review. The numbers trace back, the workbook recalculates, the slide is editable.”

— Englander, Model ML cofounder