According to Beating, Google is developing the Frozen v2 AI inference chip, which will embed part of Gemini's architecture directly into hardware to reduce computational and data movement overhead during model execution. The chip is expected to process up to 6 to 10 times more tokens per watt than Google's latest TPU, with deployment planned for as early as 2028.
The initiative aims to address Google's worsening chip shortage, which has already forced Google Cloud to decline external customer orders. Frozen v2 will operate alongside TPU; while TPU supports various models, Frozen v2 prioritizes efficiency over flexibility.