by

Google cooks up a TPU split for inference and training

Google is trying to cut Nvidia out of the AI chip gravy train.

The Alphabet unit has built a new processor tailored for the grind of querying AI models, not teaching them, as inference demand surges with businesses piling into AI agents. Google plans to unveil its eighth generation of tensor processing units, or TPUs, this week at an event in Las Vegas.

Google has been working on an inference-specific chip for several years, with people familiar with the matter saying it recently brought in a small group of AI firms to put the silicon through its paces. Alongside it, Google is rolling out a separate chip customised for training.

Google Cloud, chief executive, Thomas Kurian said: “If you don’t have inference, you cannot cover the cost of your training. So eventually inference is going to be at least as big, if not bigger, than the training market.”

That stance turns the rivalry between Google’s custom chip outfit and market leader Nvidia up another notch, as agentic AI shifts the shopping list towards faster kit that burns less power. Nvidia last month unveiled an inference play of its own, a server pairing its graphics processing units, or GPUs, with chips from the startup Groq, using technology Nvidia licensed in a $20 billion deal last year.

Another inference-leaning outfit, Cerebras, has struck a major deal with Amazon Web Services to provide AI companies with faster computing to run models. Cerebras has filed to go public, with plans to launch an IPO as soon as mid-May.

Nvidia became the world’s largest publicly listed company by flogging millions of GPUs to customers such as OpenAI, Microsoft and Oracle. The hardware is fast and brutal, performing quadrillions of simple tasks in parallel, making it handy for training.

Inference is a different beast, needing less raw compute than GPUs usually provide but more memory to keep responses snappy. Without enough memory, customers hammering lots of queries or agents can hit the “memory wall”, when data access slows and users sit there waiting.

Google Cloud, vice-president of AI and computing infrastructure, Mark Lohmeyer said: “What the customers really care about is how you can drive down the latency.”

Lohmeyer said a single request to an AI agent can generate 20 to 50 times as many “inference transactions” as a chatbot query, since the agent takes multiple actions rather than just producing text.

Google has been designing its own chips for more than a decade, with the TPU starting life inside data-centre cloud servers. As AI spread, TPUs ended up training and running Google’s generative models, including Gemini chatbots and the Nano Banana image-generator.

Each TPU generation has been built by Broadcom, which specialises in custom silicon, and Google started selling its seventh generation last year, known as Ironwood. The company has not been shy about showing off racks of TPU kit either, with one display featuring TPU 8i units glowing blue behind a tangle of colourful cables.

Last week, Google and Broadcom announced an expanded partnership to design more AI chips for Anthropic, without saying whether those parts are meant for training, inference or both.

Industry watchers have spent years wondering whether Google will properly commercialise TPUs beyond its own cloud. So far, it has cut two big outsider deals, one giving Claude-maker Anthropic access to about one million chips and another with Meta Platforms.

 

TOPICS:
ai accelerators  ·  ai agents  ·  broadcom  ·  cerebras  ·  google cloud  ·  google tpu  ·  groq  ·  inference chips  ·  nvidia rivalry

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews