OpenAI cuts GPT-5.6 Luna price 80% in full-stack AI push
On July 31, 2026, OpenAI published a post titled "Building abundant intelligence," written by Sarah Friar, laying out the company's full-stack strategy spanning infrastructure, models, platform and products. The post frames "abundance" as the economic engine behind OpenAI's business: as the cost of useful intelligence falls, more work becomes worth doing, and growing adoption feeds revenue and demand signals back into research and infrastructure investment.
As a concrete example of that cycle, the post cites the previous day's pricing announcement. GPT-5.6 Luna dropped 80% in price and GPT-5.6 Terra dropped 20%. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens; Terra costs $2 and $12 per million tokens, respectively. For GPT-5.6 Sol, Fast mode delivers up to 2.5 times the speed of standard processing at twice the price, with no change in the model's intelligence.
The post also describes efficiency work behind those price moves. GPT-5.6 Sol helped optimize the production software that serves OpenAI's models, cutting end-to-end serving costs by 20%, and helped improve speculative decoding, raising token-generation efficiency by more than 15%. Separately, improvements to retained reasoning and context management lifted GPT-5.6 Sol's score on the public ARC-AGI-3 benchmark from 13.3% to 38.3% while using six times fewer output tokens, with the underlying model itself unchanged.
On adoption, OpenAI says its models now reach more than one billion active users and more than two million businesses. Six months after signing up, people send roughly 50% more messages per day and use ChatGPT for about twice as many kinds of work, a shift the post describes as ChatGPT moving from "asking" to "doing." Across OpenAI itself, agentic work through Codex now accounts for 99.8% of weekly output tokens, with the Finance team named as one that has made agentic tools a primary part of how it operates.
On capital allocation, Friar writes that investment decisions rest on evidence: user and workload growth, enterprise commitments, API consumption, utilization, revenue, and progress in model capability and efficiency, with long-term partnerships supplying financing, infrastructure and operating expertise. She frames the goal not as more compute, bigger models or lower token prices for their own sake, but as making more useful intelligence available within reach.
Key facts
- GPT-5.6 Luna price cut 80% to $0.20 per million input tokens and $1.20 per million output tokens; GPT-5.6 Terra cut 20% to $2 and $12 per million tokens
- GPT-5.6 Sol Fast mode runs up to 2.5 times faster than standard processing at twice the price, with no change in intelligence
- GPT-5.6 Sol helped cut end-to-end model-serving costs 20% and raised speculative-decoding token efficiency by more than 15%; on the ARC-AGI-3 benchmark its score rose from 13.3% to 38.3% using six times fewer output tokens
- OpenAI's models reach more than one billion active users and more than two million businesses; users send about 50% more messages and do about twice as many kinds of work six months after signup
- Agentic work through Codex now accounts for 99.8% of OpenAI's own weekly output tokens, with Finance cited as a heavy adopter
Why it matters
The post is OpenAI's own framing of why it builds across infrastructure, models, platform and products rather than specializing in one layer: each layer feeds the others. Product use shows where customers hit friction, which shapes research; research lowers serving costs; lower costs fund more capacity. The GPT-5.6 Luna and Terra price cuts and the Sol Fast mode option are presented as the visible result of that loop, not a one-off promotion.
Who it affects
Anyone building on OpenAI's models feels the pricing shift directly: Luna and Terra get materially cheaper per token, while Sol's Fast mode offers a faster but pricier option for latency-sensitive work. The adoption figures point to a broader shift too, from individual ChatGPT users doing more with the product over time to ChatGPT Work customers and OpenAI's own teams, including Finance, running increasingly agentic workflows through Codex.
How to use it
Developers and businesses on the API can now get GPT-5.6 Luna output at $1.20 per million tokens (input $0.20) and Terra at $12 per million tokens (input $2), an 80% and 20% cut respectively from prior pricing. Workloads that need speed over cost can switch GPT-5.6 Sol to Fast mode, which runs up to 2.5 times faster at double the standard price with the same output quality; the post does not specify whether that price multiplier applies to input tokens, output tokens, or both.
How solid is it
This is a first-party strategy essay from OpenAI's own blog, authored by Sarah Friar, and every figure in it is self-reported. The specific numbers, serving-cost cuts, benchmark gains, and the pricing changes, trace back to OpenAI's own prior announcements that the post links to and summarizes rather than independently disclosing for the first time here. No outside party is cited to confirm the adoption or efficiency claims.
Risks and caveats
The post gives no baseline figures for the 'roughly 50 percent more messages' or 'twice as many kinds of work' comparisons, so the scale of that growth cannot be checked. It names no specific customers, industries, or revenue figures behind the adoption growth it describes, and gives no dollar amounts, timelines, or named partners for the 'long-term partnerships' or infrastructure investment it references. Sarah Friar's role at OpenAI is not stated in the post itself, only that she is its author.
“Better intelligence drives broader adoption. Broader adoption supports more investment. More investment improves intelligence and efficiency. That is the cycle we are building.”
— Sarah Friar, OpenAI