OpenAI cut the price of its GPT-5.6 Luna model by 80% on July 30, 2026, just three weeks after the full public launch of the GPT-5.6 family. Terra also dropped by 20%. The flagship Sol tier was left unchanged.
For enterprise buyers, the timing is the story. No model family gets repriced this aggressively this fast without a reason. Anthropic launched Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, delivering near-Fable 5 performance for roughly the same cost as its predecessor. Google shipped Gemini 3.6 Flash at $1.50/$7.50 per million tokens with a 17% efficiency improvement the same month. OpenAI’s move the following week signals that the frontier model price war has entered a new phase.
The Numbers
Luna dropped from $1/$6 to $0.20/$1.20 per million input/output tokens. That is an 80% reduction on both input and output pricing, making Luna one of the cheapest capable models available from a tier-one provider.
Terra moved from $2.50/$15 to $2/$12 per million tokens, a 20% cut. Sol remains at $5/$30 per million tokens, which positions it directly against Claude Opus 5 ($5/$25) and Gemini 3.6 Flash at the performance tier where real enterprise workloads run.
The repricing also flows through to ChatGPT Work and Codex, where usage counting now reflects the lower Luna and Terra costs. Teams running high-volume workflows through those products will see their effective bills drop.
Fast Mode Replaces Priority Processing
Alongside the price cuts, OpenAI is adding a Fast mode to the API for GPT-5.6 Sol. It delivers approximately 2.5 times faster throughput than standard processing at twice the per-token price, replacing the previous Priority Processing offering. Fast mode matters for latency-sensitive applications: real-time voice pipelines, interactive coding assistants, customer-facing chat, and any workflow where waiting on model response directly costs user experience.
The removal of Priority Processing and its replacement with the cleaner Fast Mode branding is a small signal that OpenAI is simplifying its pricing architecture as the product line matures.
Why This Matters Now
The GPT-5.6 family launched July 9 with Sol as its workhorse, Terra as its mid-tier, and Luna as its cost-efficient option. Three weeks is a very short time to revisit that pricing. It suggests the initial Luna and Terra prices did not land where OpenAI expected relative to the alternatives customers were evaluating.
Frontier model pricing has been collapsing across the board since early 2025, driven by a combination of inference efficiency gains, open-weight model improvements, and deliberate competitive moves. OpenAI’s internal narrative, shared publicly, is that Sol’s optimised inference stack is what funds the Luna and Terra cuts: better efficiency at the top of the range creates margin to compete more aggressively in the tiers where volume actually lives.
For enterprise buyers, the practical result is that the cost argument for staying on an older or less capable model gets weaker with every round of repricing. Luna at $0.20/$1.20 per million tokens is capable enough to handle a wide range of classification, extraction, and routing tasks that would have cost multiples of that six months ago.
What This Means for Business
High-volume workloads just got significantly cheaper. Any application running tens of millions of tokens per month through Luna, such as document processing, email triage, customer query classification, or batch data enrichment, will see material cost reductions without changing a line of code. The same capability costs 80% less than it did two days ago.
The competitive pressure will keep coming. OpenAI’s quick reprice is a response to the market, not a one-off event. Anthropic, Google, Meta, and a growing set of providers are all competing on the cost-performance curve. Businesses locked into a single provider at a specific price point should be building their integrations to swap models easily, because the economics will keep shifting.
Sol vs. Opus 5 is now a genuine choice. Both are priced at $5 per million input tokens. Sol has a longer track record at the performance tier and a broader ecosystem integration footprint. Opus 5 launched a week ago with notable improvements in agentic reliability and a 1-million-token context window. For businesses evaluating which to use for production deployments, both options are now real contenders at the same price.
Fast mode changes the latency calculus for Sol. Before this update, getting faster Sol throughput required workarounds or dedicated agreements. The new Fast mode gives API customers a clean way to pay for speed when their use case demands it, without changing providers or models.
For businesses currently running AI workloads, the main action is simple: check what model tier you’re on and whether the new pricing changes the build-vs-buy maths for any workflow you have been watching but not acting on yet. At Luna’s new price, there are tasks that could not justify AI processing three weeks ago that now can.
Enterprise DNA builds AI agents and custom applications that take advantage of exactly this kind of model pricing shift. If you want to understand which model tiers fit which parts of your operations, start with a discovery call.
Source
OpenAI