Jul 30, 2026
AI

OpenAI GPT-5.6 pricing cuts Luna API costs by 80%

OpenAI lowered GPT-5.6 Luna and Terra API prices while adding a higher-priced Sol Fast mode for latency-sensitive workloads.

Wei-Lin Zhao

By Wei-Lin Zhao · AI Correspondent

· 3 min read

OpenAI GPT-5.6 pricing changed sharply this week, with the company cutting Luna API rates by 80% and Terra rates by 20% while introducing a more expensive Fast mode for Sol. The move lowers the cost of running OpenAI’s smallest GPT-5.6 model in high-volume production systems and puts more pressure on Google, Anthropic and lower-cost model providers competing for agent and coding workloads.

OpenAI co-founder and CEO Sam Altman described the update on X as “major price cuts today.” OpenAI did not announce a new model generation with the pricing change, which makes the update a commercial repositioning of models that were released only recently.

How much does GPT-5.6 Luna cost now?

OpenAI says GPT-5.6 Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, or $1.40 when those two published rates are added together. Luna previously cost $1 per million input tokens and $6 per million output tokens, for a combined $7.

Terra now costs $2 per million input tokens and $12 per million output tokens, or $14 combined. Its prior combined rate was $17.50.

Sol Standard is unchanged at $5 per million input tokens and $30 per million output tokens, for a combined $35. The new Sol Fast mode costs $10 per million input tokens and $60 per million output tokens, or $70 combined. OpenAI says Fast mode can deliver up to 2.5 times the throughput without changing the model’s underlying intelligence.

Where the new prices sit against competitors

Luna is now priced below Google’s Gemini 3.5 Flash-Lite, which Google lists at $0.30 per million input tokens and $2.50 per million output tokens, or $2.80 combined. It also sits well below Gemini 3.6 Flash, priced at $1.50 per million input tokens and $7.50 per million output tokens, or $9 combined.

Luna is still not the cheapest commercial model on token price alone. Xiaomi’s MiMo-V2.5 Flash is listed at a combined $0.40 per million input and output tokens, while DeepSeek’s v4 Flash is listed at $0.42 and DeepSeek v4 Pro at $1.305. Luna’s new $1.40 combined rate puts an OpenAI frontier-series model closer to that low-cost tier, rather than the premium pricing bracket it occupied before.

Terra’s new $14 combined rate matches Google’s Gemini 3.1 Pro Preview price for context windows of 200,000 tokens or less, according to Google’s published pricing. It also creates a wider gap inside OpenAI’s own lineup: Luna costs one-tenth of Terra on a simple combined input-plus-output basis, while Terra costs 60% less than Sol Standard.

Why model pricing is tightening now

The cuts follow recent moves from OpenAI’s largest rivals. Anthropic released Claude Opus 5 at the same listed price as Opus 4.8: $5 per million input tokens and $25 per million output tokens, or $30 combined. Anthropic says Opus 5 delivers nearly all the intelligence of its more expensive Fable 5 model at about half the cost, and added an adjustable effort setting that lets developers trade reasoning depth for speed and token use.

Google, meanwhile, positioned Gemini 3.6 Flash and Gemini 3.5 Flash-Lite around lower inference costs, faster execution and agent workloads. Google has said Gemini 3.6 Flash uses fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with larger savings on some long-horizon engineering tasks.

Third-party benchmarks are now part of the pricing discussion. Artificial Analysis ranks OpenAI’s GPT-5.6 models strongly on intelligence measures, and Cognition said on X that GPT-5.6 now sits on the price-performance efficiency curve. Those are external assessments, not guarantees for a buyer’s own workload, where tool calls, context size, retries and latency requirements can change the real bill.

For enterprise teams, the competitive shift is toward total cost per completed task rather than API sticker price alone. OpenAI is lowering per-token rates for Luna and Terra, Google is emphasizing lower token use and fewer tool calls, and Anthropic is keeping Opus pricing steady while claiming better capability at the same rate. The common target is production economics for agents, coding systems, document workflows and internal assistants that can generate large token volumes quickly.

This story draws on original reporting from VentureBeat.

More from AI

All AI →