Google is exploring a server AI chip, internally referred to as Frozen v2, with the aim of directly integrating some elements of the Gemini model into hardware to improve inference efficiency. Following this news, Alphabet's stock surged over 3% intraday on July 20 but closed with a narrower gain.
The key takeaway from this news is not that Google is developing another custom chip. Investors are more concerned about whether, after the increasing cost of AI services, Alphabet can reduce the electricity, computing power, and data center pressure per question not just by continuing to ramp up capital spending, but by binding models with hardware.
The reported goals are quite ambitious. Frozen v2 aims to increase the number of tokens per unit of power service, roughly 6 to 10 times beyond the reported current or latest TPU. Tokens can be loosely understood as the basic unit for AI processing text. This figure is still a target of an undisclosed project, not a Google-confirmed mass-produced result.
The Cost Curve Imagination of Stock Price Trading
The market is pricing in the cost curve, not short-term performance, but long-term cost options. The most sensitive issue for AI investors now is whether the capital expenditure of big model companies will continue to rise and if cloud business gross margins will keep being eroded by inference costs.
Google is facing tangible pressure. The same report mentioned that the project aims to alleviate AI compute capacity shortages, as Google Cloud had rejected some external contracts due to capacity constraints. Whether these capacity constraints directly led to Frozen v2 or not, they indicate a shift: compute power is no longer just a growth input but is also becoming a constraint on revenue realization.
When ordinary users query Gemini, it consumes inference compute power. The larger the model, the more users, and the longer the responses, the more chips, power, and cooling the data center needs. For investors, this will ultimately impact capital expenditures, cloud service margins, and model service pricing.
The market implication of Frozen v2 is not just about "Google quickly solving the hash rate shortage." It is more like the long-term technical roadmap that Alphabet has outlined: if each watt of electricity can serve more tokens, Google has the opportunity to handle more AI requests under the same power and data center constraints.
Embedding Models into the Circuit Saves on Transport and Scheduling
The core of Frozen v2 is hardcoding. It can be understood as follows: a general-purpose chip is like a multipurpose kitchen that can cook any dish. Frozen v2 is more like directly writing part of Gemini's recipe into the stove, reducing the steps of looking up the recipe, moving tools, and temporarily adjusting the process each time.
AI inference is not just about computation but also involves a lot of data transport and scheduling. During model runtime, the chip needs to continuously read parameters, arrange computation paths, and move data between different storage layers. These actions consume power and time in themselves and also occupy chip resources.
If certain model structures are stable enough, they can be made into fixed circuits. The benefit is shorter paths, less scheduling, and lower power consumption. The cost is also direct – once a chip is designed based on a particular model structure, it is not easy to adapt to various new models like GPUs or TPUs.
So, the 6 to 10 times efficiency target needs to be broken down. It refers to the number of tokens that can be served per unit of power, not necessarily a 6 to 10 times reduction in Google's future data center total cost. Real chips, real workloads, and real deployment scales will all redefine the financial implications of this number.
The timeline also needs to be scaled back. Frozen v2 is reportedly not expected to be deployed until 2028 at the earliest, and project details are still being finalized. Before that, its impact on Alphabet's valuation is more akin to a long-term option rather than a cost improvement that can already be reflected in the profit and loss statement.
Custom Chips Challenge the General Narrative but Do Not Equate to Replacement
One interesting aspect of Frozen v2 is that it brings the contradiction of the AI chip roadmap to the forefront: it is difficult to maximize both universality and extreme efficiency. Over the past few years, GPUs and TPUs have excelled in adapting to different models and workloads, especially during the rapid model iteration phase.
Google's exploration appears to be more radical. Since Gemini is its flagship model, can a more tailored inference chip be designed specifically for it? As AI inference transitions from demos to everyday services, efficiency gains are no longer just engineering metrics but part of the business model.
This will pose a marginal challenge to the NVIDIA narrative, but it is not yet a "NVIDIA crisis." NVIDIA's moat still lies in the general ecosystem, software toolchains, and training clusters. If Frozen v2 holds value, it is more within specific Gemini inference scenarios at Google, reducing the marginal reliance on general-purpose computing power.
It also does not mean that the TPU is being marginalized. By all accounts, Frozen v2 is a new branch outside of the TPU, more like an efficiency tool tailor-made for the Gemini family. The TPU will still handle a wider range of training and inference tasks, serving different models and cloud customers.
A more accurate assessment is that when a model is sufficiently important, has a large enough number of calls, and has a stable enough architecture, dedicated chips begin to make economic sense. It is not a signal of the industry's full-scale shift to dedicated chips, but rather a sign that leading model companies are starting to separately account for high-frequency inference scenarios.
Gemini's Stability Determines the Value of This Option
The biggest risk of Frozen v2 lies not only in whether the hardware team can create a more efficient circuit, but also in whether the underlying architecture of Gemini can stabilize enough to justify being hardcoded into the chip. AI models are still evolving rapidly, and what is optimal today may not be the best solution two years from now.
If Gemini undergoes significant architectural changes in the future, Frozen v2 may need to be redesigned. This would extend the timeline and weaken the cost advantage brought by hardcoding. The deeper the specialization, the greater the efficiency potential, but the higher the cost of changing course once the technical path is locked in.
What the market needs to validate is whether this path can progress from early project hints to a more solid engineering milestone. The official roadmap, tapeout progress, mass production timeline, actual energy efficiency metrics, and deployment scale will all determine whether this will be a new anchor on Alphabet's cost curve or a forward-looking experiment whose valuation has already been priced in.
For Alphabet, the current significance of Frozen v2 is to show the market that the pressure of AI spending may not necessarily have to be addressed solely by continuing to increase capital expenditure. However, until 2028, it is more like an option. What can truly change the valuation is not just the code name itself, but whether this chip can turn Gemini's high-frequency calls into measurable cost reductions.
Welcome to join the official BlockBeats community:
Telegram Subscription Group: https://t.me/theblockbeats
Telegram Discussion Group: https://t.me/BlockBeats_App
Official Twitter Account: https://twitter.com/BlockBeatsAsia
