OpenAI launches GPT-6.1 Sol, near GPT-6 Astra at one-fifth the price

OpenAI launches GPT-6.1 Sol, near GPT-6 Astra at one-fifth the price

OpenAI has introduced GPT-6.1 Sol, an upgrade to GPT-6 Sol. The company says it nearly matches GPT-6 Astra's intelligence on agentic coding, computer use and professional work, at one-fifth of Astra's standard input and output token prices. Cached input costs $0.10 per million tokens, which is 95% less than GPT-6.1 Sol's standard input price and 50% less than GPT-6 Sol's cached input price. OpenAI frames this as more room for developers to build agents that reuse context across requests.

The pitch is a new balance of capability and cost for everyday work: writing and debugging code, understanding documents and running multi-step business workflows. OpenAI claims substantial gains over GPT-6 Sol on complex professional tasks, and says that on several evaluations the new model approaches, but does not match, Astra's performance at substantially lower cost.

The benchmark results are mostly given as margins and cost ratios rather than absolute scores:

  • DeepSWE v1.1 (complex software-engineering tasks in real codebases): GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost, and beats GPT-6 Sol's best score by 6.4 percentage points at a lower reasoning effort and cost.
  • GDP.pdf (answering professional questions from complex PDFs with tables, charts, diagrams and fine print): it scores higher than Opus 5.5 with fallbacks at less than half the cost per task across the tested reasoning settings, and approaches Astra's state-of-the-art performance at roughly one-fifth the cost per task.
  • AutomationBench (whether agents correctly complete multi-step business workflows): 2.2 percentage points above Opus 5.5 at medium reasoning effort, at roughly a third of the cost, and up 4.8 percentage points from GPT-6 Sol at the same setting.
  • OSWorld 2.0, offline set (demanding computer-use workflows): it outperforms GPT-6 Sol by seven percentage points at maximum reasoning effort at less than half the cost, and comes within 2.1 percentage points of Astra's score at maximum effort at roughly one-seventh the cost per task.
  • Terminal-Bench Science 0.1 (scientific workflows including data analysis, simulation and theorem proving): it more than doubles GPT-6 Sol's score at maximum reasoning effort at less than half the cost per task. At maximum effort it costs $5.47 per task on average, against $23.21 for Opus 5.5 and $23.80 for Astra, which OpenAI describes as over 75% lower cost than either model. Astra still posts the highest score among the models tested, 68.1%, and OpenAI says it should be used for the most difficult scientific research tasks.

On factual accuracy, the largest improvement over GPT-6 Sol comes at low reasoning effort, where the share of responses containing a factual error falls from 11.4% to 7.7%, a reduction of approximately 32%. Across the tested reasoning settings, GPT-6.1 Sol's error rate stays within 1.9 percentage points of Astra's, at less than one-fifth the cost per task. The test counts answers with at least one factual error on de-identified conversations where users flagged an earlier model's mistake. OpenAI notes these deliberately difficult prompts are not representative of typical usage.

On alignment, OpenAI reports substantial improvements over GPT-6 Sol, bringing the model closer to Astra. It says GPT-6.1 Sol is more transparent about its limitations and more reliable at respecting user intent and safety constraints. In challenging evaluations it shows lower failure rates than GPT-6 Sol on transparency about broken search tools, respecting explicit restrictions and avoiding unauthorized outcomes during agentic tasks. OpenAI observed no attempts to bypass an automated safety reviewer, matching GPT-6 Astra and GPT-6 Sol. Full details are in a system card addendum. These evaluations deliberately test challenging situations and do not measure failure rates in typical use.

Availability: from the day of the announcement, GPT-6.1 Sol is available to all Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex. It is not yet available in Chat. Developers can call it through the OpenAI API as gpt-6.1-sol, at $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. In the coming days OpenAI will also offer GPT-6.1 Sol Ultrafast, with up to 8x faster token generation than standard speed in Codex.

OpenAI states that evaluations of its own models were run in its research environment or via its API, which may give slightly different output from production ChatGPT because of differences in system prompts, tools and effort settings. Evaluations of competitor models were taken from publicly available reports.

Key facts

  • GPT-6.1 Sol, an upgrade to GPT-6 Sol, is said to nearly match GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra's standard input and output token prices.
  • API model name is gpt-6.1-sol: $2 per million input tokens, $0.10 per million cached input tokens, $10 per million output tokens.
  • On Terminal-Bench Science 0.1 at maximum effort it costs $5.47 per task on average, versus $23.21 for Opus 5.5 and $23.80 for Astra; Astra still leads on score at 68.1%.
  • Share of responses with a factual error on difficult prompts drops from 11.4% to 7.7% at low reasoning effort versus GPT-6 Sol, about 32% fewer.
  • Available now to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, not yet in Chat; an Ultrafast variant with up to 8x faster generation in Codex is due in the coming days.

Why it matters

The announcement is about price-performance. OpenAI says GPT-6.1 Sol closes most of the gap to its top model, GPT-6 Astra, on agentic coding, computer use and professional work while costing about a fifth as much per token. On DeepSWE v1.1 it claims to match Astra at roughly one-fifth of the cost. It also claims to beat Opus 5.5 on GDP.pdf and AutomationBench at a fraction of the cost per task. The 95% cheaper cached input, and 50% below GPT-6 Sol's cached price, targets agents that reuse context across many requests.

Who it affects

Developers building agents on the API get a cheaper model priced at $2, $0.10 cached and $10 per million tokens. Plus, Pro, Business, Enterprise and Edu users get it in ChatGPT Work and Codex. Teams currently paying for Astra or Opus 5.5 on coding, document-heavy, business-workflow or computer-use tasks are the obvious audience for the cost comparisons. Researchers doing the hardest scientific work are pointed back to Astra.

How to use it

In ChatGPT Work and Codex the model is available from the day of the announcement for Plus, Pro, Business, Enterprise and Edu users. It is not yet in Chat. Via the API, call gpt-6.1-sol. Standard prices are $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. Structuring prompts so context is reused across requests is where the cached rate helps most. GPT-6.1 Sol Ultrafast, with up to 8x faster token generation than standard speed in Codex, is promised in the coming days; no price or exact date is given. For the most difficult scientific research tasks, OpenAI says to use GPT-6 Astra.

How solid is it

This is OpenAI's own announcement, and every result comes from the company. Its own-model evaluations were run in its research environment or via its API, which can differ slightly from production ChatGPT. Competitor numbers were taken from publicly available reports. No independent third-party verification is mentioned. Most results appear as margins in percentage points or cost ratios, not absolute scores; the one absolute score given is Astra's 68.1% on Terminal-Bench Science 0.1. Details of the alignment work sit in a separate system card addendum.

Risks and caveats

"Nearly matches" is not "matches": OpenAI says the model approaches Astra on several evaluations, and Astra still has the top Terminal-Bench Science 0.1 score and is recommended for the hardest scientific tasks. The factuality test uses deliberately difficult prompts that OpenAI says are not representative of typical usage, and the error rate stays within 1.9 percentage points of Astra's, not below it. The alignment evaluations likewise test challenging situations and do not measure failure rates in typical use. The $5.47 per-task figure is stated as an average, and the source does not say whether it is a mean or a median. GPT-6.1 Sol is not yet in Chat, and Ultrafast is still to come.

“GPT‑6 Astra still achieves the highest score among the models tested at 68.1%, and should be used for the most difficult scientific research tasks.”

— OpenAI, GPT-6.1 Sol announcement