AMD reckons its Ryzen AI Halo workstation can save local AI coders from cloud bills, provided they fancy a $3,999 punt.
The Ryzen AI Halo, AMD’s answer to Nvidia’s DGX Spark AI workstation, will be available for pre-order next month. It starts at $3,999, which is still cheaper than a midlife crisis.
AMD claims the box could save developers $750 a month if they spend eight hours a day vibe coding with local models instead of cloud APIs. That is the sales pitch, although anyone doing eight hours of vibe coding daily may have bigger problems.
The price looks punchy for an AI mini PC. Less than a year ago, similar hardware could be found for between $2,200 and $2,999, but the RAMpocalypse has apparently made everyone suffer.
Nvidia’s DGX Spark now retails for $4,699, up from $3,999 when it was reviewed last autumn. AMD’s version is trying the same trick by offering a curated developer environment for running local models and agentic AI frameworks.
These machines are not the fastest AI boxes on the block. Their real job is letting developers run models that, only a few years ago, would have needed systems costing $20,000 or more.
The Ryzen AI Halo measures 5.9 x 5.9 x 1.7 inches, or 150 x 150 x 43mm. Inside sits a 120 watt Ryzen AI Max+ 395 APU, better known as Strix Halo. The chip gets 128GB of LPDDR5x 8000MT/s memory feeding 16 Zen 5 cores and 40 RDNA 3.5 GPU compute units. AMD says that provides up to 256GB/s of bandwidth, which is more than a Ryzen 9000 Threadripper non-Pro system.
For local AI types, that is enough to run models up to 200 billion parameters at 4-bit precision. That matches the more expensive DGX Spark, which will annoy Nvidia’s marketing elves.
Most of the compute comes from the integrated graphics, which can deliver about 56 teraFLOPS at 16-bit precision. That sounds impressive until Nvidia’s Blackwell-based GB10 APU wanders in with bigger numbers.
The DGX Spark offers 125 teraFLOPS at BF16, 250 at FP8 and 500 at FP4. Double those figures if a workload can use Nvidia’s 2:4 sparsity.
AMD’s Strix Halo does not support FP8 or FP4 data types in hardware. That leaves the Ryzen AI Halo between 55 and 88 per cent slower than Spark on advertised floating point performance.
That gap will not always show up in real work. AMD claims the AI Halo generates tokens 4-14 per cent faster than the Spark in LLM inference. A similarly equipped HP Z2 Mini G1a, using the same silicon, reportedly edged ahead of the Spark in Llama.cpp with the Vulkan backend. That suggests AMD’s little box is not just a spec-sheet ornament.
Token generation is mostly dictated by effective memory bandwidth rather than raw floating point swagger. GPU compute matters more for prompt processing, where Nvidia’s tensor cores have the advantage.
In testing cited by the source, the Spark had a 2x to 3x lead in prompt processing. Short prompts only added tiny delays, but longer prompts made the gap harder to ignore. The Spark was ahead in image generation and fine-tuning benchmarks too. AMD’s software stack has improved since then, so the gap may have narrowed a bit.
AMD has two decent cards to play. The Ryzen AI Halo includes an XDNA 2-based NPU rated at 50 TOPS, though its usefulness depends on the application. Many content creation apps can use the NPU, but generative AI inference engines have been less keen. That may change, assuming developers enjoy another layer of compatibility fun.
The second advantage is that the Ryzen AI Halo is a standard x86 box. Users can run Windows or their favoured Linux flavour, instead of being nudged into one blessed lane. Nvidia’s Spark runs a lightly customised Ubuntu 24.04 setup.
For developers building for Microsoft’s NPU-accelerated AI PC ecosystem, AMD’s approach is the obvious fit. Networking is less flattering. Nvidia’s AI workstation includes a 200Gbps ConnectX-7 NIC for clustering two systems and eventually four.
AMD’s AI Halo gets a single 10Gbps NIC. That is fine for hauling large model files, but it is not the same league for clustering. High-speed networking over USB-4 might be possible, although support is unclear. The Fruity Cargo Cult Apple has shown RDMA over Thunderbolt, so AMD might yet have a playbook somewhere in the drawer.
Much of the value comes from validated hardware, documented playbooks and known-good software. That matters because AI developers have long suffered through dependency soup made from ROCm, HIP, SYCL, CUDA, PyTorch, TensorFlow and JAX.
At launch, AMD says the Ryzen AI Halo will ship with five preinstalled playbooks. Another 10 will be available online, with more added monthly. Customers will get access to AMD’s developer programme, cloud credits and exclusive playbooks. That should reduce the time spent swearing at drivers, although it will not remove the sport entirely.
The 128GB Ryzen AI Halo will be available for pre-order next month starting at $3,999. AMD is already preparing a 192GB version for people who think 128GB is for peasants. That higher-capacity system will use the refreshed Ryzen AI Max+ 495 APU. Like the rest of AMD’s 400-series lineup, it gets modest CPU, GPU and NPU clock bumps rather than architectural fireworks.
The 192GB unified memory option should open the door to larger models. The price will probably open the door to an awkward conversation with whoever approves your expenses.







