by

AMD and Intel smoke peace pipe to fix x86 AI

AMD and Intel have published the full specification for AI Compute Extensions, better known as ACE, setting up a new x86 standard for AI matrix work on future processors.
The specification, now at version 1.15, gives software developers a stable target for AI and high-performance computing libraries before compatible chips arrive.
That last part is worth underlining, since there is no announced ACE-capable silicon yet. Hardware is not expected until around 2028, so the standard has arrived before the machines that can actually use it.
AI workloads lean heavily on matrix multiplication, and traditional x86 SIMD extensions were never a natural fit for that job. AVX and its successors are built around vector processing. That means they handle long one-dimensional streams of values, while AI matrix work is a two-dimensional problem. For years, CPUs have been forcing the wrong tool to do the job, because apparently even silicon enjoys office politics.
GPUs solved this years ago with dedicated tensor hardware. CPUs, by comparison, have been useful for many AI tasks only when there was no better option available.
ACE changes the structure by adding eight two-dimensional tile registers to x86. Each tile can store a 16-by-16 matrix of 32-bit values, giving the CPU a more natural way to chew through matrix operations.
The new instructions use an outer-product approach, allowing much more matrix work to be done per instruction than a comparable AVX10 operation.
AMD and Intel claim this gives ACE up to 16 times the matrix-compute density of an equivalent AVX10 multiply-accumulate operation with the same number of input vectors.
That does not mean every AI workload will suddenly run 16 times faster, because real performance will depend on memory bandwidth, compiler support and how much die space each chipmaker gives to ACE hardware.
Still, the instruction overhead reduction is real. Fewer instructions are needed to do the same matrix work, which should help efficiency and reduce the pointless shuffling that makes CPUs look faintly embarrassed next to GPUs.
ACE supports AI-friendly data formats including INT8, FP8, BF16 and block-scaled formats from the Open Compute Project. That gives it a more modern AI profile than AVX10.
Intel already has Advanced Matrix Extensions, or AMX, in Xeon server chips, but AMD and Intel have chosen not to turn AMX into the shared standard.
Instead, ACE is a fresh extension backed by both companies. Eight of the named authors on the whitepaper are AMD engineers, with three from Intel, which says plenty about how the balance of power in x86 has shifted.
AMX was built for Intel’s server world. ACE is meant to cover a wider range of x86 devices, from servers to laptops and embedded systems, depending on how each vendor implements it.
ACE will not make x86 CPUs competitive with Nvidia GPUs for heavy AI training or the largest inference workloads. Its purpose is to give x86 a common AI compute baseline for the workloads that already run on CPUs, but currently do so inefficiently.
If AMD and Intel ship ACE consistently, x86 could finally gain a unified matrix compute path that software developers trust.
The chipmakers have not made CPUs into GPUs, but they have at least admitted that pretending vector extensions were enough for AI was getting silly.

TOPICS:
ace  ·  ai compute extensions  ·  AI inference  ·  AMD  ·  avx10  ·  Intel  ·  intel amx  ·  matrix multiplication  ·  X86

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews