Jul 31, 2026
AI

Deepseek V4 Flash 0731 closes in on OpenAI Luna at lower cost

Deepseek’s updated budget model scored 50 on Artificial Analysis, one point below GPT-5.6 Luna, with lower reported task costs.

Renata Fuchs

By Renata Fuchs · Policy Reporter

· 3 min read

Deepseek V4 Flash 0731 closes in on OpenAI Luna at lower cost
Photo: The Decoder

Deepseek V4 Flash 0731 has been released as an upgraded version of the company’s budget AI model, with Artificial Analysis scoring it at 50 on its Intelligence Index. That is 10 points above the prior V4 Flash model from April 2026 and one point below OpenAI’s GPT-5.6 Luna, according to Artificial Analysis.

The commercial signal is pricing. Artificial Analysis said the new Deepseek model costs about 60 percent less per task than GPT-5.6 Luna, even after OpenAI reduced pricing on that model by 80 percent. Deepseek did not disclose a new funding round, revenue figure or customer count in the materials cited, so the measurable news is benchmark position, published model characteristics and price-to-performance claims.

How does Deepseek V4 Flash compare with GPT-5.6 Luna?

Artificial Analysis places Deepseek V4 Flash 0731 almost level with OpenAI’s budget Luna model on its broad index, while giving Deepseek the stronger price-to-performance position. The reported score gap is narrow: 50 for Deepseek’s updated model versus 51 for GPT-5.6 Luna.

Deepseek’s pricing advantage appears to come partly from its cache discount. Deepseek’s API documentation lists a 98 percent discount for cached usage, compared with what Artificial Analysis describes as a 90 percent industry-standard discount. A cache discount lowers the cost when repeated context can be reused rather than processed at full input-token pricing.

The new model also uses 12 percent fewer tokens than its predecessor, according to Artificial Analysis. For operators running high-volume inference workloads, token efficiency matters because output quality alone does not determine production cost. A model that reaches comparable benchmark scores while consuming fewer tokens can change the deployment math for customer support, coding agents and internal workflow tools, assuming the benchmark gains translate into real usage.

What changed in the updated model?

Artificial Analysis said Deepseek V4 Flash 0731 improved over the prior version in every tested category. The largest gains were in agentic tasks, where models are evaluated on their ability to carry out multi-step work rather than answer isolated prompts.

On GDPval, a benchmark intended to measure performance on complex office-style knowledge work, the model rose from 1,189 to 1,559 Elo points, according to Artificial Analysis. The firm also reported that the updated model hallucinates less often, though the cited materials do not provide a user-facing error rate or production reliability figure.

The underlying architecture did not change. Deepseek V4 Flash 0731 keeps 284 billion total parameters, with 13 billion active parameters, and supports a one-million-token context window. Those specifications keep it in the same efficiency-oriented design category rather than signaling a new frontier-scale architecture.

The weights for Deepseek V4 Flash 0731 are available on Hugging Face under an MIT license, according to the model listing. That licensing detail matters for teams that want to inspect, host or modify the model rather than rely only on a hosted API, although deployment cost will still depend on infrastructure, workload shape and latency requirements.

The release keeps pressure on budget AI pricing. OpenAI had already cut GPT-5.6 Luna pricing sharply, but Artificial Analysis’ comparison suggests Deepseek is still using cache economics and token efficiency to compete below OpenAI’s cost per task while staying close on benchmark scores.

This story draws on original reporting from The Decoder.

More from AI

All AI →