by

Nvidia’s CUDA moat keeps AI rivals coughing dust

Nvidia’s CUDA software stack keeping the world’s machine-learning work welded to its GPUs rather than any of its hardware.

Wired’s Sheon Han argues that CUDA, not “a piece of hardware,” is the outfit’s proper moat and why rivals can wave decent specs around and still look second-best.

“What sounds like a chemical compound banned by the FDA may be the one true moat in AI,” Han wrote.

CUDA’s main trick is parallelisation. Instead of making one processor grind through a 9×9 multiplication table one sum at a time, a GPU can split the work across cores.

A nine-core GPU can hand one column to each core and get a ninefold speed gain. With more optimisation, it can recognise that 7×9 equals 9×7 and avoid doing the same job twice.

That cuts 81 operations to 45, which matters when a single training run can cost $100 million. At that point, shaving off wasted work is not polish, it is survival.

Nvidia’s GPUs began life rendering graphics for games. In the early 2000s, Stanford PhD student Ian Buck saw that their architecture could be used for general high-performance computing.

Buck had come to GPUs through gaming, created a programming language called Brook, then joined Nvidia. With John Nickolls, Buck helped lead CUDA’s development.

“If AI ushers in the age of a permanent white-collar underclass and autonomous weapons, just know that it would all be because someone somewhere playing Doom thought a demon’s scrotum should jiggle at 60 frames per second,” Han wrote.

CUDA has become a packed bundle of software libraries for AI, with each function shaving nanoseconds from maths operations. Those tiny cuts add up.

Han puts it like this: “Let’s say the task is peeling garlic. An unoptimized GPU would go: “Peel the skin with your fingernails.” CUDA can instruct: “Smash the clove with the flat of a knife.” PTX lets you dictate every sub-instruction: “Lift the blade 2.35 inches above the cutting board, make it parallel to the clove’s equator, and strike downward with your palm at a force of 36.2 newtons.”

That sort of control is why Nvidia’s software lead is so hard to copy. It is not enough to build a chip with tasty benchmark numbers and hope developers arrive carrying flowers.

“You can begin to see why CUDA is so valuable to Nvidia — and so hard for anyone else to touch. Tuning GPU performance is a gnarly problem. You can’t just conscript some tender-footed undergrad on Market Street, hand them a Claude Max plan, and expect them to hack GPU kernels. Writing at this level is a grindsome enterprise — unless you’re a cracker-jack programmer at DeepSeek…”

Han says rivals such as AMD and Intel can offer competitive specifications on paper. Chipzilla and AMD, however, still have to deal with software stacks that have struggled with bugs, compatibility problems and weak adoption.

TOPICS:
ai  ·  AMD  ·  CUDA  ·  deepseek  ·  GPUs  ·  Intel  ·  machine learning  ·  Nvidia  ·  PTX

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews