Google Is Reportedly Baking Gemini's Architecture Directly Into Silicon With 'Frozen v2'

HardwareLLMs
Illustration generated by AI: Editorial image for Google Is Reportedly Baking Gemini's Architecture Directly Into Silicon With 'Frozen v2'

The Core · TL;DR

  • Google is reportedly developing 'Frozen v2,' an internal server chip that embeds Gemini's model architecture directly into silicon rather than model weights.
  • The chip could be 6 to 10 times more power-efficient than Google's current AI chips, measured by tokens generated per watt, but isn't expected until 2028.
  • The design concept originated with Jeff Dean, Google DeepMind's chief scientist, and the chip is meant for internal compute capacity, not external sale.
  • Google shares rose about 3% after The Information's report, as the company plans $180-190 billion in AI infrastructure spending.

Google is working on a server chip so specialized that it hard-wires the structure of its Gemini models into the silicon itself, according to a report from The Information. The project, internally known as Frozen v2, is designed to squeeze dramatically more performance out of every watt Google spends running its AI systems, and the company reportedly expects it to be between six and ten times more efficient than its current AI chips when measured by tokens generated per unit of power.

What sets Frozen v2 apart from a typical accelerator is what gets fixed at the hardware level. Rather than embedding a specific set of trained model weights, the chip encodes Gemini's underlying architecture, the fundamental design choices that govern how the model processes information. That distinction matters: weights change every time a model is retrained or fine-tuned, but an architecture tends to persist across generations of a model family. Locking that structure into silicon lets Google strip away the flexibility that general-purpose chips need, trading it for raw efficiency on the specific computations Gemini relies on.

The concept reportedly traces back to Jeff Dean, Google DeepMind's chief scientist, whose original "Frozen" design has now evolved into this second-generation chip. Unlike Google's TPU line, which is offered to outside cloud customers, Frozen v2 is being built strictly for internal use. The goal isn't to sell a new product but to relieve pressure on Google's own AI compute capacity as demand for running Gemini at scale continues to climb.

That framing lines up with the sheer scale of Google's infrastructure spending. The company has earmarked between $180 billion and $190 billion for AI infrastructure, a figure that underscores just how much strain serving billions of Gemini queries is putting on its data centers. A chip purpose-built to cut power costs per token, even one arriving years from now, represents a meaningful lever against that spending curve.

Timing is the biggest caveat here. Frozen v2 isn't expected to ship until 2028, meaning any efficiency gains are still roughly two years off and could shift as the project moves through development. Markets, however, reacted quickly to the report itself: Google's stock rose about 3% the Monday morning the story broke, suggesting investors are already pricing in the prospect of Google reducing its own inference costs at a moment when AI compute expenses are under intense scrutiny across the industry.

If Google can deliver anywhere close to the projected efficiency jump, the implications extend beyond its balance sheet. Architecture-specific silicon could become a template other hyperscalers follow as they look for ways to keep serving increasingly capable models without a proportional increase in power consumption and data center footprint.

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram

More from Research

View all in Research