Google reportedly developing 'Frozen v2' chip with Gemini's architecture etched into the silicon — engineers project 6 to 10 times more tokens per watt than latest TPUs

Google
(Image credit: Google)

Google is developing a server chip, informally dubbed "Frozen v2," that would etch part of its Gemini model's architecture directly into the silicon, according to a report published Monday by The Information, citing two people with direct knowledge of the matter. Engineers on the project have projected that the chip could serve six to ten times more tokens per unit of power than the newest generation of Google's TPUs, with deployment targeted for as soon as 2028. The two sources said the project is partly a response to an AI compute shortage severe enough that Google Cloud has turned down deals with outside customers.

Latest Videos FromTom's Hardware
TOPICS
Luke James
Contributor

Luke James is a freelance writer and journalist.  Although his background is in legal, he has a personal interest in all things tech, especially hardware and microelectronics, and anything regulatory. 

  • usertests
    Model-hardwired inference silicon already exists in demonstrations. Taalas, a Toronto startup that has raised more than $200 million from investors including Quiet Capital and Fidelity, launched its HC1 chip in February with Llama 3.1 8B permanently wired into an 815mm-squared die built on TSMC's N6 process. The company claims 17,000 tokens per second per user with no HBM on the package.
    The Taalas chip apparently uses 250W. If something like that was in a PCIe card, it could be interesting for the home user.
    Reply
  • MonkoftheFunk
    Bingo, came to say the same thing. Seeing how fast Chatjimmy is, this would be awesome for open claw, robots or real time translation. I really think this would pop the bubble if consumers got this.
    Reply
  • KryptonQuark
    Awesome News ! ! ! !
    Reply
  • usertests
    Model-hardwired inference silicon already exists in demonstrations. Taalas, a Toronto startup that has raised more than $200 million from investors including Quiet Capital and Fidelity, launched its HC1 chip in February with Llama 3.1 8B permanently wired into an 815mm-squared die built on TSMC's N6 process. The company claims 17,000 tokens per second per user with no HBM on the package. Nvidia, meanwhile, struck a $20 billion deal in December to license technology from inference chip designer Groq.
    AMD has acquired Taalas:
    https://wccftech.com/amd-snaps-up-taalas-weeks-after-cerebras-deal-chasing-chips-that-bake-ai-models-into-silicon/https://newsroom.amd.com/news/amd-acquires-taalas-ai-inference/
    Reply
  • usertests
    3MKRjt59hh4
    AMD's AI accelerators could end up offloading expensive agentic loops to Taalas-designed chips.

    I hope we see this for consumers.
    Reply