Alphabet is developing a new server chip, internally dubbed "Frozen v2," designed to run its Gemini models up to ten times more efficiently than the company's current AI chips, according to a report by The Information cited across multiple outlets1,2.
The chip would permanently embed parts of Gemini's architecture directly into the silicon, reducing the number of calculations and amount of data movement required to answer queries. Google engineers project Frozen v2 could serve between six and ten times more tokens per unit of power than Google's newest TPUs. Deployment is targeted for 20283.
Alphabet shares climbed 3% on Monday following the report. In a statement to CNBC and TechCrunch, Google did not directly confirm or deny the project but said its teams are "constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers" and that "while not every project moves into production, this rigorous exploration is central to our full stack approach". The company added that "by co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized for real-world workloads".
ANALYSIS The approach of baking model-specific logic into silicon represents a departure from the general-purpose flexibility of standard TPUs, trading versatility for raw inference efficiency on a single model family.
The chip effort comes against a backdrop of internal compute pressure. Google's cloud unit has reportedly been forced to turn away outside business because of a chip shortage. Google is reportedly paying SpaceX nearly a billion dollars a month for additional capacity.
TechCrunch noted that AI companies have increasingly sought to produce their own chips to make in-house models run more efficiently and to address global shortages in AI computing capacity, while also attempting to reduce dependence on Nvidia, which has historically dominated the AI chip market.
ANALYSIS If the projected efficiency gains hold at scale, Frozen v2 could materially lower Alphabet's per-query inference costs for Gemini workloads — a metric that matters as efficiency becomes a competitive differentiator among AI providers.
The report lands with Alphabet's earnings two days away, with the stock on pace for a third straight monthly decline. Investors are expected to look for proof that the company's record AI spending is translating into faster cloud growth and a stronger backlog.
Separately, Google DeepMind CEO Demis Hassabis is meeting with lawmakers in Washington, D.C., pitching a federally overseen, industry-funded body to test advanced AI models before release. That lobbying effort coincides with the departure of the Commerce Department's top official in charge of AI testing after just three months on the job.
Google also faces execution questions beyond the chip timeline: the next Gemini Pro release is reportedly delayed, senior researchers have left the lab, and Chinese rivals are narrowing the capability gap, according to CNBC's reporting.