OpenAI cuts GPT-5.6 Luna pricing by 80% as AI model costs fall
OpenAI lowered GPT-5.6 Luna to $0.20 per million input tokens and $1.20 per million output tokens, intensifying AI price pressure.
By Renata Fuchs · Policy Reporter
· 3 min read
OpenAI cut GPT-5.6 Luna pricing by 80% effective July 30, taking its lowest-cost GPT-5.6 model to $0.20 per million input tokens and $1.20 per million output tokens. The move matters for AI buyers because inference costs remain one of the main constraints on deploying generative AI at scale, and OpenAI is signaling it will compete harder on unit economics, not only frontier performance.
The company also reduced pricing for GPT-5.6 Terra by 20%, bringing that model to $2 per million input tokens and $12 per million output tokens. Pricing for GPT-5.6 Sol is unchanged, according to OpenAI. The company said all three models are available through ChatGPT Work, Codex and the OpenAI API.
What is GPT-5.6 Luna pricing now?
GPT-5.6 Luna now costs $0.20 per million input tokens and $1.20 per million output tokens. OpenAI said Luna delivers performance comparable to leading models from a year earlier, while a workload that previously cost $1 on those models now costs about 6 cents on Luna and runs nearly nine times faster.
Those are company claims, and OpenAI did not disclose customer volume, revenue impact or gross margin assumptions tied to the new prices. The practical implication for operators is clearer: if Luna is good enough for a workload, the new pricing changes the threshold for tasks such as coding assistance, summarization, classification and other high-volume inference jobs where cost per token can determine whether a feature ships.
Why did OpenAI lower GPT-5.6 prices?
OpenAI said the reductions were enabled by efficiency gains from GPT-5.6 Sol inside its own infrastructure. In a company post, OpenAI said Sol optimized GPU software on its own, reducing deployment costs by 20%. The company also said token generation improved by more than 15% through speculative decoding, a technique that speeds output by using a smaller model to draft tokens before a larger model verifies them.
The price cut also lands during a broader squeeze on AI model pricing. Low-cost Chinese AI providers have pushed aggressive price-to-performance comparisons, and Microsoft has begun promoting its own MAI models as lower-cost alternatives to OpenAI. That competitive pressure gives enterprise buyers more leverage, especially for workloads that do not require the most capable model in a provider’s lineup.
For frontier labs, the downside is that cheaper inference can pressure revenue growth unless usage expands enough to offset lower unit prices. That is a central issue for companies carrying large infrastructure commitments: model providers need to keep improving performance while also filling expensive compute capacity. OpenAI’s Luna cut suggests the company is willing to trade price per token for adoption, at least on its most affordable GPT-5.6 tier.
The more durable question for the market is whether model pricing keeps falling faster than customer demand rises. OpenAI’s announcement gives developers an immediate cost reduction on Luna and Terra, but it does not answer how much of the savings can be sustained if the AI price war continues.
This story draws on original reporting from The Decoder.