Aug 13, 2026
AI

Gemini 3.7 Flash pricing starts at half rate for coding and agents

Google’s Gemini 3.7 Flash launches with temporary API rates 50% below standard pricing and vendor-reported gains in coding and workflow tests.

Renata Fuchs

By Renata Fuchs · Policy Reporter

· 3 min read

Google announced Gemini 3.7 Flash on Aug. 13, positioning the model for coding, AI agents and knowledge work, with Gemini 3.7 Flash pricing set at $0.75 per million input tokens and $3.75 per million output tokens through Dec. 31, 2026. The rates are half Google’s stated standard prices, which rise to $1.50 for input and $7.50 for output on Jan. 1, 2027.

The release came about three weeks after Gemini 3.6 Flash, which Google launched July 21 at the same $1.50 and $7.50 standard rates. For teams running high-volume coding or document-processing agents, the offer lowers the initial API bill, but it is a time-limited discount rather than a permanent reduction.

Context caching is priced at $0.075 per million tokens during the introductory period and $0.15 per million tokens from Jan. 1, 2027, according to Google’s published pricing cited in reporting on the launch.

How much does Gemini 3.7 Flash cost?

  • Through Dec. 31, 2026: $0.75 per million input tokens, $3.75 per million output tokens and $0.075 per million cached tokens.
  • From Jan. 1, 2027: $1.50 per million input tokens, $7.50 per million output tokens and $0.15 per million cached tokens.

Those rates are only one part of an agent’s operating cost. A deployment also consumes tokens through reasoning and tool calls, and may incur retries, infrastructure charges and human-review work. Google says 3.7 Flash makes more deliberate plans, handles roadblocks better and follows instructions more closely, with the aim of cutting retries and manual oversight. Those are company claims, not independently established production results.

Google reports gains in coding and enterprise workflows

In Google-reported tests, 3.7 Flash scored 43.6% on FrontierCode 1.1 Main, compared with 34.4% for 3.6 Flash. Its DeepSWE v1.1 result was 65.3%, up from 49.0%, while its Code Arena score reached 1,588 Elo from 1,538 for the prior version.

Google also reported improvements outside software development: 30.4% on AutomationBench, versus 17.0% for 3.6 Flash, and 34.0% on GDP.PDF, compared with 22.0%. The benchmarks point to the workloads Google is targeting, including multi-step automation and document comprehension, rather than establishing a universal model leader.

Google’s own comparison table shows trade-offs. GPT-5.6 Terra scored 69.6% on DeepSWE v1.1, ahead of 3.7 Flash’s 65.3%, and it led the listed Terminal-bench 2.1, Terminal-bench 3.0 and OSWorld-2.0 comparisons. Claude Sonnet 5 led the cited Agent’s Last Exam multimodal desktop and operating-system task, at 33.3% versus 26.3% for 3.7 Flash.

Operators should therefore test representative tasks rather than extrapolate from a leaderboard. A useful evaluation tracks completed tasks, total tokens, tool calls, retries and the rate of human corrections. Venture Post’s guide on evaluating AI models for the work they will actually do outlines why those measures matter alongside benchmark scores.

Where is Gemini 3.7 Flash available?

Google AI Pro and Ultra subscribers can use 3.7 Flash in Gemini Spark, Google’s personal AI agent. Google says Spark can use the model for Workspace-related work such as consolidating files, drafting emails and updating status documents. The model is also available through the Gemini Enterprise Agent Platform and the Gemini Enterprise app.

Google did not provide a release date for Gemini 3.5 Pro with the 3.7 Flash announcement. The company had previously said that model was being tested with partners and would be broadly released when ready.

This story draws on original reporting from VentureBeat.

More from AI

All AI →