Drooid Logo
Back to story perspectives

Full Breakdown

Google Targets AI Efficiency with “Frozen v2” Server Chip

7/20/2026, 8:36:35 PM

Core Development: “Frozen v2” Chip for Gemini Models

Alphabet’s Google is engineering a new server-side processor, internally codenamed Frozen v2, that will embed core elements of its Gemini large-language model directly into silicon. By hard-wiring the model’s architecture, the chip is projected to handle six to ten times more tokens per unit of power than Google’s latest Tensor Processing Units (TPUs). The design is a specialized branch of Google’s custom-chip portfolio and is not intended to replace the general-purpose TPUs. Sources say the first deployment could occur as early as 2028.

Background & Context

Google’s internal AI teams have been grappling with a severe compute capacity crunch that has sparked tensions across engineering groups and forced Google Cloud to turn away some external customers. The company’s TPUs, originally built for internal search workloads, have become a cornerstone of its cloud AI services, but the surge in generative-AI demand has outpaced existing hardware. Across the industry, rivals such as OpenAI, Anthropic, Nvidia, Amazon, Microsoft, Meta and Apple are also pursuing purpose-built silicon to gain performance and cost advantages.

Data & Statistics

  • Efficiency gain: 6–10 × tokens per power unit versus current TPUs.
  • Share reaction: Alphabet stock rose 3 % to $359.68 in early trading after the report.
  • Internal pressure: Google Cloud has declined certain external deals due to the compute bottleneck.

Official Statements & Responses

Google’s communications emphasize ongoing research and a “full-stack” approach, noting that not every exploratory project reaches production. The company frames Frozen v2 as a means to deliver higher performance and efficiency for users while maintaining flexibility through “flexible hard-wiring” of model architecture rather than fixed weights.

Criticism & Opposition

Analysts caution that embedding a specific model architecture limits future adaptability; if Gemini’s underlying design changes, the chip could become obsolete. The “rigidity” trade-off is viewed as a gamble, especially as competitors ship models that may surpass Gemini’s capabilities. Internal reports also highlight that the project heightens resource-allocation tensions within Google’s AI divisions.

On-the-Ground Reports

Employees familiar with the effort describe a “compute bottleneck” that has forced Google Cloud to reject some partnership proposals. A separate Bloomberg report noted a delay in the launch of Gemini 3.5 Pro after it fell short of internal coding benchmarks, underscoring the pressure to improve hardware efficiency.

Conflicting Reports & Gaps

  • Deployment date: Some sources say “as soon as 2028,” while others cite a “targeted early-2028” timeline.
  • Scale of production: Reports indicate Frozen v2 will not be produced at the same volume as TPUs, but exact manufacturing plans remain undisclosed.
  • Official confirmation: Google has not formally confirmed the chip’s specifications or release schedule.

Verbatim Quotes

  • “Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers,” — Google Cloud spokesperson
  • “By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized for real-world workloads.” — Google statement
  • “while not every project moves into production, this rigorous exploration is central to our full stack approach.” — Google statement
  • “not every project reaches production, but this rigorous exploration is core to our full-stack approach.” — Google statement

Why It Matters

If Frozen v2 achieves its efficiency targets, Google could lower operating costs, accelerate Gemini-driven services, and strengthen its position in the escalating AI-hardware race. The chip’s success would illustrate the strategic shift toward tightly coupled hardware-software stacks, potentially reshaping how large-scale AI models are deployed across cloud platforms.