by

Cerebras gives AMD a smug moment

AMD’s Cerebras has picked up an OpenAI compliment, giving Helios a shine no amount of roadmap fairy dust could buy.

According to Reuters, OpenAI researcher Jeffrey Wang praised Cerebras chips for inference speed.

“Internally, we have some OpenAI models that are on Cerebras chips. They’re incredible because they have such fast inference,” Wang said.

For those not in the know, Helios is AMD’s full-stack rack-level AI box, built around Instinct MI455X GPUs, sixth-generation EPYC CPUs, Pensando networking, Infinity Fabric and ROCm.

AMD wants to sell the whole AI rack, not just another accelerator. Cerebras brings its Wafer-Scale Engine, a giant piece of interconnected silicon with enough on-chip memory and compute.

Its trick is to keep much of the model work close to SRAM, rather than sending data on a scenic tour through network bottlenecks.

AMD wants Helios to handle throughput, while Cerebras handles low-latency decoding and token generation, with the pair claiming up to five times more tokens per second per watt.

“And what this means for me on the day-to-day is: whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive,” Wang said.

That matters because slow inference turns clever AI into a coffee break with syntax highlighting.

If the answer lands before a developer wanders into email or Slack, the silicon has done its job and the human has less time to forget why they started.

Reuters reported on 23 July 2026 that OpenAI plans to start using Helios racks later this year. AMD chief executive Lisa Su said Helios is in full production, with shipments slated to start near the end of the third quarter.

 

TOPICS:
AI inference  ·  AMD  ·  cerebras  ·  helios  ·  Instinct MI455X  ·  Nvidia  ·  openai  ·  rocm  ·  Wafer-Scale Engine

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews