by

Who needs an NPU GPUs already do the job?

Chip vendors are flogging NPUs as the future of AI PCs, while your mid-range gaming laptop is doing four times the maths.

The marketing hook is simple: buy a Co-Pilot Plus PC because its NPU can hit 45 TOPS, which sounds impressive until you notice an RTX 4060 can push more than 200 TOPS without shouting about it.

Writing for XDA-Developers, Jasmine Mannan points out that if you want to run an LLM locally or generate images on your machine, the GPU is still the undisputed king.

So why are Intel, AMD and Qualcomm leaning so hard on NPUs if a discrete GPU cleans the clock with them on raw throughput?

The answer is less about peak AI bragging rights and more about battery life, thermals and a dose of platform control.

Mannan said that an NPU is designed for always-on, low-power AI tasks such as webcam background blur, eye-contact correction or noise suppression, typically sipping around 5W.

A discrete GPU is a brute-force parallel monster that can pull 50 to 100W, finish the job far faster and then sit back, having taken a noticeable bite out of your battery.

If you need to generate 50 images, the NPU will eventually get there, but the GPU will do it in a fraction of the time.

That difference matters if you are on the move, because a thin-and-light with an NPU can run AI-enhanced features all day without sounding like a jet engine.

It matters far less if your laptop lives on a desk most of the time, plugged in and pretending to be a desktop replacement. Mannan also flags software maturity as a key gap.

Nvidia’s CUDA and AMD’s ROCm have had years to mature, so most local AI tools such as Stable Diffusion, LM Studio and Ollama work out of the box on a GPU.

NPUs, by contrast, are fragmented and often require specific runtimes, such as OpenVINO on Intel or ONNX on Windows, and many open-source projects provide only limited support.

Then there is memory, which in 2026 is not cheap, as AI data centre demand is pushing up RAM prices.

An NPU shares system RAM, so if you have 16GB, it is competing with your browser and the OS for model storage.

A GPU has its own VRAM, whether GDDR6 or GDDR7, with far higher bandwidth than DDR5, which makes a tangible difference in tokens per second when running a local model such as Llama 3.

That means a 32GB gaming laptop with a decent GPU can be a better AI machine than a new AI-branded laptop stuck at 16GB with an NPU.

Vendors push NPUs because they are efficient, integrate seamlessly into the SoC, and enable badge machines as AI-ready without adding a bulky graphics card.

For light AI features and long battery life, that makes sense.

For creators, developers, or enthusiasts who want to run models locally, a gaming laptop with a suitable GPU remains the more capable option, regardless of how many TOPS the marketing slide claims.

 

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews