Jul 21, 2026
AI

Google is reportedly designing a Gemini-specific AI inference chip

The Information says Google’s Frozen v2 chip could improve inference efficiency by 6 to 10 times, with deployment planned for 2028.

Renata Fuchs

By Renata Fuchs · Policy Reporter

· 3 min read

Google is reportedly designing a Gemini-specific AI inference chip
Photo: The Decoder

Google is developing a server chip known internally as Frozen v2 that would hardwire parts of its Gemini model architecture into silicon, according to The Information. The reported goal is lower-cost AI inference: sources cited by The Information said the chip could be 6 to 10 times more efficient at serving AI responses than Google’s current TPU chips.

The chip is not expected to be a broad replacement for Google’s TPUs. The Information reported that Google plans to start deploying Frozen v2 in 2028 and views it as an experiment in more specialized AI chips, with production volumes below those of its TPU line. Google has not disclosed cost targets, expected unit volumes, manufacturing partners or whether the chip would be used across all Gemini traffic.

The design marks a different bet from the general-purpose accelerator strategy used by most cloud AI infrastructure. Google’s TPUs are built to support multiple models and are now part of its external cloud pitch. Frozen v2, as described by The Information, would trade some flexibility for efficiency by placing elements of Gemini’s architecture directly into hardware.

Why the design is different

The reported chip name refers to the practice of “freezing” parts of a model so they no longer change. In this case, the frozen component would be silicon-level support for part of the model structure, rather than a software setting. That could reduce the amount of computation required to generate responses and improve latency or cost per query, depending on how much of the architecture is fixed in hardware.

The Information reported that the original concept came from Jeff Dean, chief scientist at Google DeepMind. His first design would have embedded model weights directly into a chip. Weights are the learned parameters that shape a model’s responses. Google dropped that approach, according to the report, because it would have tied the chip to a single version of Gemini and risked becoming obsolete as the model changed.

Frozen v2 is described as a more flexible version of that idea. Instead of locking in weights, it would embed the model’s underlying architecture while allowing new weights to be loaded. The Information reported that Google has not yet decided how much of Gemini’s architecture will be hardcoded.

Inference margins are the point

The commercial logic is straightforward. Training frontier models is expensive, but inference costs determine the economics of running those models at scale. Any reduction in the cost of generating responses can affect margins, product pricing and capacity planning, especially for companies serving consumer and enterprise AI products at high volume.

Frozen v2 also sits apart from Google’s TPU commercialization push. Google leases TPUs to Meta, offers TPUs to Google Cloud customers and has positioned its TPU@Premises program as an alternative to Nvidia systems. The Information and The Decoder have reported that Google has an internal target of capturing 10% of Nvidia’s annual revenue with TPUs. Frozen v2, by contrast, appears aimed at Google’s own compute needs rather than external customers.

That limitation follows from the design. A chip built around Gemini’s architecture would be useful only as long as Google keeps that architecture stable enough to justify the silicon. That makes it less attractive as a cloud product for customers running a range of models, but potentially valuable for Google if Gemini traffic becomes large and predictable.

If the reported efficiency gains hold, Frozen v2 could help Google run Gemini at lower cost than rivals using less specialized infrastructure. The Information did not report whether Google has working silicon, how the chip compares with future TPU generations, or whether the 2028 deployment target is fixed.

This story draws on original reporting from The Decoder.

More from AI

All AI →