by

Microsoft rolls out Maia 200 inference chip

Software King of the World, Microsoft, wants everyone to know it has a new inference chip and it thinks the maths finally works.

Volish executive vice president Cloud + AI Scott Guthrie introduced Maia 200 as a first-party accelerator aimed at cutting the cost of AI token generation, not winning a beauty contest.

Under the heatsink, Microsoft says Maia 200 is built on TSMC’s 3nm process, with native FP8 and FP4 tensor cores, 216GB of HBM3e pushing 7TB/s, plus 272MB of on-chip SRAM.

Vole claims Maia 200 is its most efficient inference system so far, with 30 per cent better performance per dollar than the latest kit it is already running.

The company is pitching Maia 200 as part of a mixed hardware estate, slated to serve multiple models, including OpenAI’s GPT-5.2, and it says the Microsoft Superintelligence team will use it for synthetic data generation and reinforcement learning.

Microsoft’s line is that Maia 200’s design accelerates synthetic data pipelines, generating and filtering domain-specific data more quickly, so downstream training gets fresher signals rather than stale leftovers.

Maia 200 is already running in Vole’s US Central data centre region near Des Moines, Iowa, with US West 3 near Phoenix, Arizona, next, and other regions after that.

Developers get a Maia SDK preview, according to Microsoft, with PyTorch integration, a Triton compiler, an optimised kernel library and access to a low-level language, which sounds like the usual “here are the tools, please port your stack” dance.

For raw compute, Microsoft is quoting more than 140 billion transistors per chip and output of more than 10 petaFLOPS at 4-bit precision and more than five petaFLOPS at 8-bit, inside a 750W SoC TDP envelope.

Microsoft dedicates a fair chunk of its press release to data feeding, saying that the chip has a redesigned memory subsystem centred on narrow-precision data types, a specialised DMA engine, on-die SRAM and an on-chip fabric for high-bandwidth movement.

At the system level, Maia 200 uses a two-tier scale-up network built on standard Ethernet, with Microsoft pointing to a custom transport layer and integrated NICs rather than proprietary fabrics.

Each accelerator exposes 2.8TB/s of bidirectional scale-up bandwidth and can run collective operations across clusters of up to 6,144 accelerators, with four accelerators per tray linked directly to keep chatty traffic local.

Vole said it validated the backend network and a second-generation closed-loop liquid-cooling Heat Exchanger Unit early, and wired in telemetry and management down to the chip and rack.

The most eye-catching brag is the schedule, with Microsoft claiming models were running on Maia 200 within days of the first packaged parts arriving, and that the time from first silicon to first rack deployment was cut to less than half that of comparable programmes.

It takes swings at rivals by claiming three times the FP4 performance of third-generation Amazon Trainium, and FP8 performance above Google’s seventh-generation TPU, which is the sort of benchmark talk that always depends on the small print.

 

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews