by

CUDA starts leaking onto AMD GPUs in Windows

A solo developer has managed to run CUDA-targeted applications on an AMD Radeon RX 9060 XT natively under Windows using ZLUDA and ROCm.

According to Tom’s Hardware, the CUDA-for-AMD-Windows on GitHub, created by developer Speedstu, provides a scripted compatibility setup for software stubbornly written with Nvidia hardware in mind.

Rather than creating another GPU runtime, the project integrates ZLUDA’s CUDA translation layer with AMD’s HIP and ROCm stack on Windows. The result lets CUDA-facing applications talk to a Radeon without virtualisation, WSL2 or a Linux dual boot.

That could be interesting for developers stuck with older repositories, proprietary applications or AI software whose authors apparently concluded the computing world begins and ends with Nvidia CUDA.

AMD has improved official Windows support for PyTorch and its HIP SDK, including Radeon RX 7000 and RX 9000 hardware. CUDA-only applications remain another matter.

Speedstu’s scripts detect the AMD GPU and its gfxXXXX architecture, check the HIP installation and download a pinned official ZLUDA build. The validated configuration uses ZLUDA v6-preview.69, AMD HIP SDK 6.4 and LibTorch 2.3.0+cu118. So far, only the Radeon RX 9060 XT using AMD’s gfx1200 RDNA4 target has been validated. Other Radeon GPUs are treated as candidates rather than supported hardware, and the developer is asking users to report successful and failed tests.

The project intercepts CUDA driver calls before passing them through ZLUDA and into AMD’s libraries. CUDA’s cuBLAS calls can be handled through rocBLAS, while cuBLASLt maps to hipBLASLt and cuSPARSE uses rocSPARSE. cuFFT also works on the tested setup.

The CUDA driver interface, nvcuda, passed the project’s runtime checks alongside cuBLAS, cuBLASLt, cuSPARSE and cuFFT. Speedstu demonstrated the setup using a real 2,216,347-parameter PPO reinforcement-learning network compiled against CUDA-enabled LibTorch.

The network completed forward and inference operations, PPO learning and optimiser work on the Radeon, with one validation run processing 65,536 timesteps. A controlled test on 13 September 2026 ran 10 iterations of the same workload on the RX 9060 XT.

After throwing away the first warm-up iteration, the public upstream configuration managed a median throughput of 13,278 steps per second. A recovered custom overlay from earlier development reached 12,876 steps per second, leaving the clean upstream configuration about three per cent quicker.

However, Speedstu noted that “a later rewrite removed LibTorch/ZLUDA from PPO and achieved substantially higher throughput,” showing there is still a performance cost from piling translation layers between the application and hardware. There are also some fairly substantial holes.

The stable Windows HIP SDK used for testing does not provide cuDNN support through this setup. Software relying heavily on Nvidia’s deep neural network library is therefore likely to hit a wall.

NCCL, TensorRT, unsupported PTX behaviour and some custom CUDA extensions can fail and Windows still exposes only part of the broader ROCm ecosystem available elsewhere.

“This does not mean every CUDA program or AI model works. CUDA API/library coverage is workload-dependent,” the project documentation said.

ZLUDA itself is hardly backed by an army of engineers. The project lost commercial backing and is now maintained as what its developer calls a “weekend hobby project.”

That makes betting a production environment on it a courageous move in the traditional corporate sense. Still, the Windows demonstration attacks an awkward part of Nvidia’s software advantage. CUDA’s enormous software ecosystem has long made changing GPU suppliers more complicated than swapping one graphics card for another. Software written specifically for CUDA can keep users tied to Nvidia even when competing hardware can do the adding up just as well.

CUDA-for-AMD-Windows does not make that problem disappear, but an RX 9060 XT running an unmodified CUDA-facing LibTorch workload shows the lock is not entirely welded shut.

 

TOPICS:
AMD  ·  CUDA  ·  gpu computing  ·  Nvidia  ·  Radeon  ·  rocm  ·  rx 9060 xt  ·  windows  ·  zluda

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews