Intel and AMD are pushing AI Compute Extensions as a unified path to make x86 great again.
Last year, Intel and AMD teamed up to bolster the x86 ecosystem through their “x86 Ecosystem Advisory Group” initiative, aiming to establish a standard feature set across architectures.
The promise was a cleaner baseline that stays accessible, scalable and compatible with what comes next, with four features flagged early on: FRED, AVX10, ChkTag and ACE.
Now the ACE whitepaper from AMD and Chipzilla is out, laying out what the pair think this extension can do for x86 chips.
They claim the EAG helped align and refine the ACE ISA so that matrix acceleration features land in a standard form across the ecosystem, with both vendors contributing ideas and the community providing market reach.
They say they are still cooperating on the future roadmap for ACE and AVX10, aiming to pursue “new opportunities in AI and other workload domains” without dumping developers into yet another dead end.
The paper frames ACE as a seamless add-on to AVX10, offering low-friction matrix acceleration that is meant to be everywhere in x86, not just in a few big iron parts.
ACE targets a “significant increase in matrix multiply performance” alongside scalability and energy efficiency, because matrix multiplication is the core grind inside neural networks and LLM workloads.
SIMD extensions such as AVX10 can perform matrix multiplication, yet compute density and scalability are limited, and the usual tricks for accelerating matrix multiplication are not considered efficient.
The EAG pitch is that ACE accelerates matrix multiplication with greater flexibility, letting existing AVX10 optimisation carry forward from laptops right up to supercomputers.
That cross-platform angle is meant to reduce developer pain compared with pushing AI compute out to specialist hardware every time the workload grows teeth.
In the whitepaper, AMD and Chipzilla call ACE the “Standard Matrix Acceleration Architecture for x86”, which is the sort of line that dares compiler teams to prove them wrong.
ACE support covers native matrix multiplication for common AI formats such as INT8, OCP FP8, OCP MXFP8, OCP MXINT8 and BF16, with an outer product operation designed to work with AVX10.
They claim the ACE outer product approach delivers a 16x compute-density bump over an equivalent AVX-10 multiply-accumulate operation, using the same number of input vectors.
Because ACE extends AVX10, software enablement is already underway, with integrations for deep learning and HPC libraries, plus Python staples such as NumPy and SciPy, and frameworks like PyTorch and TensorFlow.







