Google’s Strategic Pivot: Developing a Dedicated Gemini Chip to Reshape AI Computing Economics

Google is working on a new chip for serving its AI models that are optimized to deliver high performance on Google’s Gemini AI models, which could change the economics of AI deployments. The custom silicon, known as “Frozen v2,” marks a shift from traditional silicon-based hardware approaches to AI, by integrating the underlying architecture of Google’s AI powerhouse directly into the silicon itself, according to The Information’s exclusive reporting. As the company strives to streamline its AI operations, reduce costs, and meet growing demands, the project proves to be a timely and essential initiative.

The move for this specialized chip is motivated by a critical need of this tech giant due to an internal reality: Google is facing a crunch in its ability to process AI-related workloads, which industry observers say has caused tensions within the company. The shortage has reached critical proportions, forcing Google Cloud to reject potential business deals with outside customers, according to several sources familiar with the situation, a sign of the strain on the company’s AI infrastructure. The slowdown could affect Google’s own AI research as well as its plans to emerge as a major cloud-based AI provider that rivals Microsoft and Amazon, which are also rolling out their own AI services.

The report estimates Google is aiming to roll out the Frozen v2 chip as early as 2028, but cautions the timing is still up in the air as engineers are still working to complete the design architecture. One of the most important decisions yet to be made is the exact amount of Gemini’s operating system that will be baked into silicon, a determination that will balance performance improvements and flexibility. The consequences are significant, as the design choices made here today will impact the chip’s ability to keep pace with future versions of Gemini, or become quickly outdated as AI technology changes.

image

Frozen v2’s performance numbers are nothing short of phenomenal. The preliminary estimate is that the chip could be 6-10 times more efficient than Google’s custom AI chips in terms of AI tokens processed per unit of power consumption. These would be driven mostly by removing the need to process and react to different model architectures, as chips would be flooded with model variations throughout the process of interpreting and responding, rather than being tasked with performing specific operations that have been pre-optimized for a particular AI architecture. This is a basic efficiency boost that could change the economics of getting big language models to scale, enabling Google to compete with its rivals for AI services at lower cost points using the same technology that relies on traditional computing gear.

One interesting aspect of Frozen v2 is the fact that it is related to Google’s current Tensor Processing Unit (TPU) family of processors. The new chip isn’t designed to out-compete the TPU ecosystem; instead, it’s meant as a complement to it, a parallel line of chips specifically tailored to specific deployment situations. The TPU architecture continues to be essential to Google’s AI infrastructure, and the new Ironwood TPUs are now included in the eighth iteration of the architecture, which will continue to power the wide range of machine learning workloads and external customers. Frozen v2, on the other hand, is a wager on the idea that Google can be specialized enough to become more efficient, at the cost of some flexibility, but with such performance gains being critical in the fiercely competitive AI space.

The project is named “Frozen” and that’s exactly what it is: an engineering project. The name comes from the idea of “freezing” parts of the Gemini model into the physical structure of the chip, so that they can be applied without having to make decisions in real-time for the inference operation. The design is a departure from the typical AI accelerators such as Nvidia’s GPUs or even Google’s own TPUs, which are usually used for various AI frameworks and dynamic model architectures that demand considerable computational resources. Frozen v2 could significantly cut down on the delay and power usage of each AI query, making it possible for fresh sorts of apps to demand almost instantaneous responses and extremely low running costs.

The technology implications of such a solution pervade beyond just the technical and extend to the very economics of cloud AI services. Industry analysts have long highlighted that the large-scale inference cost of a large language model is a big hurdle in the path of widespread adoption, and in some cases the cost of inference may surpass the cost of training the model over its lifetime when deployed at scale. A significant cost benefit for Google in running apps with Gemini might prove to be a game-changer in the AI landscape.The efficiency gains promised by Frozen v2 could give Google a competitive edge in operating apps powered by Gemini, potentially altering the AI industry landscape. This benefit could mean cheaper customer prices, better profit margins, or entirely new AI services that simply wouldn’t be feasible using existing hardware.

👁️ 61.3K+
Kristina Roberts

Kristina Roberts

Kristina R. is a reporter and author covering a wide spectrum of stories, from celebrity and influencer culture to business, music, technology, and sports.

MORE FROM INFLUENCER UK

Newsletter

Sign up for Influencer UK news straight to your inbox!