★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Hardware

Run an LLM on Your Laptop\'s NPU: What Actually Works

October 4, 2026
16 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Got the hardware sorted? Now build on it. You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Start free
Or own it for life — Lifetime $149, pay once

Short answer: your NPU shows 0% because Ollama has no NPU backend at all — its own hardware docs list NVIDIA, AMD ROCm, Apple Metal and Vulkan, and nothing else. LM Studio does not have one either; its runtime picker only chooses between CPU and GPU engines. To actually reach the NPU you need software written for that specific silicon: FastFlowLM or Lemonade on AMD Ryzen AI (works today, closest thing to an Ollama experience), OpenVINO GenAI on Intel Core Ultra (works, but you must export the model first), and the experimental Hexagon backend in llama.cpp on Snapdragon X (build it yourself, no installer). Intel's friendlier ipex-llm route is archived and should not be used.

That is the whole answer. The rest of this page is the per-chip detail, the exact commands, the driver versions, a section on why swapping runtimes in LM Studio changes nothing, and one section on the comparison everybody wants and nobody has published.


Why Your NPU Sits at 0%

It is not a setting. Ollama, every runtime LM Studio ships and llama.cpp's mainline builds do not target NPUs, so there is nothing to enable. LM Studio is the confusing one because it does hand you a runtime chooser — that gets its own section below.

An NPU is not a small GPU. It has its own driver stack, its own memory-movement model, and it runs precompiled kernels rather than the general-purpose shader or CUDA code that GPU inference engines emit. That means a runtime cannot "also support NPU" as a checkbox — someone has to write and ship compiled kernels per operator, per quantisation format, per NPU generation. That work has happened for a handful of projects and not for the ones you already have installed.

Ollama's hardware documentation is explicit about the accelerators it supports: NVIDIA GPUs at compute capability 5.0+, AMD Radeon via ROCm, Apple Metal, and additional GPUs via Vulkan on Windows and Linux. AMD Ryzen AI parts appear in that documentation as Radeon integrated graphics, which is the iGPU — not the XDNA NPU. Intel Core Ultra NPU and Qualcomm Hexagon are absent entirely.

So when you bought a machine advertised at 50 TOPS and Task Manager's NPU graph stays flat while the CPU cooks, nothing is broken. You are simply running software that was never pointed at that block of silicon. If you want the buying-side view of what those TOPS ratings mean across vendors, that is our NPU comparison; this page is about making the thing move.

Mac owners, take this exit now. The Apple Neural Engine is the same story with a harder ceiling: it is reachable only through Core ML, there is no public low-level API to compile and dispatch kernels the way you can with Metal or CUDA, and MLX's maintainers closed the ANE request in December 2023 saying they have no plans while it stays a closed-source API. Nothing on this page will light up your ANE, because nothing does. The full chain — Ollama, llama.cpp, MLX, and the powermetrics command to watch your own Mac — is in does Ollama use the Apple Neural Engine?. Everything from here down is about Windows laptops.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

Verdict Per Chip

AMD is usable today, Intel is usable with an extra export step, Snapdragon is a build-it-yourself project.

Your chipCan you run an LLM on the NPU?SoftwareFriction
AMD Ryzen AI (XDNA2) — Strix, Strix Halo, Kraken, Gorgon PointYesFastFlowLM, or Lemonade (which uses FastFlowLM as its NPU backend)Low — Windows MSI installer, Ollama-style CLI
Intel Core Ultra Series 2 (Lunar Lake) and Core Ultra NPUsYes, with caveatsOpenVINO GenAIMedium — model must be exported to INT4 OpenVINO IR first; static-shape limits
Qualcomm Snapdragon XExperimentalllama.cpp Hexagon backendHigh — compile from source with the Hexagon SDK
Ryzen AI 300 with older XDNA1 (Phoenix/Hawk Point)Not via FastFlowLM—FastFlowLM's stated support is XDNA2 parts
Apple Silicon M-series (Neural Engine)No — and not through any runtime—Blocked at the API layer, not the silicon: ANE vs Metal, explained

One cross-cutting note before the details: on all three platforms the iGPU and CPU paths still work normally and are far less trouble. Running on the NPU is a deliberate choice you make for power or for freeing the GPU, not a prerequisite for local AI on these laptops. If you just want models running tonight, run them on the CPU and come back to this page later.


AMD Ryzen AI (XDNA2)

This is the one that works like you hoped. FastFlowLM installs from an MSI and its CLI is a deliberate Ollama clone.

FastFlowLM (1.8k stars) describes itself as Ollama built for AMD Ryzen AI NPUs. Supported silicon, per its README: all Ryzen AI series chips with XDNA2 NPUs — Strix, Strix Halo, Kraken and Gorgon Point. Windows gets an flm-setup.msi installer; Linux support arrived in March 2026. Version 1.0.0 landed in August 2026 under the headline "First Release Under ROCm" — the project now ships as part of AMD's ROCm organisation, which matters more for its longevity than any benchmark on this page. Licensing: MIT for the orchestration code and CLI, with the NPU kernels free for any use including commercial.

Requirement to check first: NPU driver 32.0.203.304 or newer (the project recommends the .311 build). Older Ryzen AI driver packages will not work. You also need internet access on first run so it can pull the compiled NPU kernels for your model.

flm list                    # models available
flm run llama3.2:1b         # interactive, on the NPU
flm serve llama3.2:1b       # OpenAI-compatible server on port 52625

Inside a session, /verbose turns on performance reporting and /bye exits — again, straight out of the Ollama playbook. Because flm serve speaks the OpenAI API on port 52625, anything you already point at a local endpoint (Open WebUI, a VS Code extension, your own script) works by changing the base URL.

Model coverage is broad for a project this young: the documented families include LLaMA, Qwen, Gemma, gpt-oss, Phi, DeepSeek, LiquidAI LFM, Whisper and EmbeddingGemma, with the docs claiming full context length support across the list. If gpt-oss is what you are after, our gpt-oss model page covers the weights themselves.

The published speed numbers

FastFlowLM publishes per-model benchmark pages. Their LLaMA results, measured on an AMD Ryzen AI 7 350 (Kraken Point) with 32GB DRAM, FastFlowLM v0.9.30, default performance mode:

ModelDecode (tok/s @ 1k context)Prefill (tok/s @ 1k prompt)
LLaMA 3.2 1B64.51,686
LLaMA 3.2 3B26.3766
LLaMA 3.1 8B12.8403

Source: FastFlowLM benchmark documentation. These are vendor-published figures on one specific chip, with no iGPU or CPU baseline alongside them — read the measure-it-yourself section before drawing conclusions.

Take the shape of that table seriously even if you discount the absolute values: prefill is where the NPU looks strong (1,686 tok/s ingesting a prompt on a 1B model is a lot for a thin-and-light), while decode at 12.8 tok/s for an 8B is readable-but-not-fast. That profile — fast at chewing through context, moderate at generating — is worth knowing when you decide what to use it for.

Lemonade: the same NPU, a wider net

Lemonade (5.4k stars, Apache-2.0) is the broader runtime: it serves LLMs across NPUs, GPUs and CPUs — XDNA2 NPU, NVIDIA GPUs from Turing on, AMD RDNA3/4 including the Strix Halo iGPU, Apple Silicon, and x86_64/ARM64 CPUs. It also handles speech-to-text, TTS and image generation, not just text. Its NPU support is FastFlowLM under the hood: the v11.0.0 release notes describe the FastFlowLM NPU backend auto-installing on Linux rather than needing a manual system package. v11.6.0 shipped 14 August 2026, so it is moving fast.

lemonade backends          # what your machine can actually use
lemonade run Gemma-4-E2B-it-GGUF

Pick Lemonade if you want one runtime that falls back gracefully across NPU, iGPU and CPU, or you want the non-text modalities. Pick FastFlowLM directly if you only care about the NPU and want the smallest possible stack. For the Strix Halo class of machine specifically, see our AI Max+ 395 guide and the best Strix Halo mini PC roundup — on those boxes the 128GB unified memory pool is a bigger story than the NPU.


Intel Core Ultra / Lunar Lake

The NPU works, but the tutorial you found is probably for a dead project. Intel archived ipex-llm on 28 January 2026.

Start with the bad news, because it will save you an evening. The repository intel/ipex-llm — 8.9k stars, the project behind nearly every "run an LLM on your Intel NPU" blog post from 2024 and 2025 — is archived. Its banner reads:

Intel will not provide or guarantee development of or support for this project, including but not limited to, maintenance, bug fixes, new releases or updates.

GitHub additionally flags it as having known security issues. Its NPU quickstarts are still visible and still indexed, which is exactly why people keep landing on them. Do not build anything new on it.

The maintained path is OpenVINO GenAI. It is more work than ollama run, but it is documented, current and supported.

# 1. Export and compress the model to OpenVINO IR, INT4 symmetric channel-wise
optimum-cli export openvino -m meta-llama/Meta-Llama-3.1-8B-Instruct \
  --weight-format int4 --sym --ratio 1.0 --group-size -1 \
  Meta-Llama-3.1-8B-Instruct

Then load that exported directory with an OpenVINO GenAI LLM pipeline and pass "NPU" as the device. Things the documentation is clear about, and that you should plan around:

  • NPU driver: install the latest; the troubleshooting guidance names 32.0.100.3104 or newer when execution fails.
  • Weight formats accepted on NPU: INT4 symmetric channel-wise, INT4 group quantisation at group size 128, NF4 (Series 2 NPUs and later only), and INT8 for Whisper models. You cannot just point it at a random GGUF — this is a different model format entirely.
  • Static shapes. The NPU pipeline compiles for fixed shapes, which is what makes it fast and also what constrains it: the documented defaults are a maximum input prompt of 1024 tokens and a minimum response length of 128 tokens.
  • RAM. Intel documents that Core Ultra Series 2 systems may need more than 16GB of RAM to process prompts longer than 1024 tokens with models above 7B parameters. On a 16GB Lunar Lake ultrabook, plan on smaller models.

That 1024-token default prompt ceiling is the honest headline for Intel. It is adjustable, but the static-shape design means long-context chat is not what this path is built for. Summarising a short email, running a local autocomplete, powering a small assistant: fine. Feeding it a 30-page PDF: use the iGPU. Intel's discrete cards are a different and much less constrained story — see Intel Arc B580 for local AI.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

Snapdragon X

There is a real Hexagon NPU backend in llama.cpp, it is labelled experimental, and there is no installer.

llama.cpp carries Snapdragon documentation under docs/backend/snapdragon covering three backends on these devices: CPU, Adreno GPU via OpenCL, and Hexagon NPU (HTP). Facts worth having before you commit a weekend:

  • Platforms: Android arm64, and native Windows 11 arm64 builds. Linux appears in the Docker toolchain instructions.
  • Hexagon versions: the docs reference v73, v75, v79 and v81, varying by hardware generation.
  • Quantisation: the worked examples use Q4_0 (e.g. Llama-3.2-1B-Instruct-Q4_0.gguf). Treat Q4_0 as the tested format.
  • Memory model: models under 4B fit in one Hexagon session; an 8B needs two; a 20B needs four. You control this with GGML_HEXAGON_NDEV.
  • How offload works: the docs note the Hexagon NPU "behaves as a GPU device when it comes to -ngl and other offload-related options", so the flags you already know carry over.
  • Debugging: GGML_HEXAGON_PROFILE and GGML_HEXAGON_VERBOSE environment variables.
  • Published speed: roughly 51.5 tok/s token generation for Llama-3.2-1B in their tg64 benchmark. That is one model at one size on unstated silicon — take it as evidence the path works, not as a spec.

The cost is the build. The Docker toolchain image bundles the Android NDK, OpenCL SDK, Hexagon SDK and CMake; a native Windows-on-Snapdragon build wants Visual Studio (for the Windows SDK and Driver Kit), LLVM/Clang, CMake, Hexagon SDK 6.6.0.0 and the Adreno OpenCL SDK 2.3.2. There is no MSI and no brew install.

Our honest recommendation for Snapdragon X owners: unless you enjoy toolchains, run models on the CPU — these are competent ARM cores and llama.cpp's CPU path on Windows-on-ARM is mature — and revisit the Hexagon backend when it loses the experimental label. Nothing about it is fake; it is just not a consumer install yet.


Does LM Studio Support the Intel or Ryzen AI NPU?

LM Studio has no NPU runtime on any platform. The list you are choosing from contains CPU engines and GPU engines, so switching runtimes moves the model between the processor and the graphics block — never onto the Intel, Ryzen AI or Snapdragon NPU.

This deserves its own answer because LM Studio is the one app that makes the mistake reasonable. Ollama hides its backend selection entirely, so nobody expects a knob. LM Studio puts a runtime chooser right in front of you, and a screen full of selectable engines invites the obvious guess: one of these must be the NPU. None of them is.

The app's own requirements are the tell. LM Studio's system requirements ask for "AVX2 instruction set support is required (for x64)" and "at least 4GB of dedicated VRAM is recommended" — one CPU requirement and one GPU requirement. Apple Silicon gets a chip line. Windows on ARM is listed as "ARM (Snapdragon X Elite) based systems". An NPU is not mentioned on any of the three platforms, and neither is a TOPS floor. A vendor that shipped NPU acceleration would say so on that page.

What each runtime actually targets on your machine

Your machineWhat the runtime list is offering youNPU reached?Where the NPU path actually lives
Intel Core Ultra (Lunar Lake, Arrow Lake)CPU builds, plus the Arc integrated graphics as a GPU deviceNoOpenVINO GenAI, a separate toolchain with its own export step — Intel section above
AMD Ryzen AI 300 / AI Max (XDNA2)CPU builds, plus the Radeon integrated graphics as a GPU deviceNo. Radeon graphics and the XDNA NPU are two different devices behind two different drivers — a graphics API cannot address the NPUFastFlowLM, or Lemonade — AMD section above
Qualcomm Snapdragon X EliteThe ARM64 CPU. The requirements page names the architecture and stops — no Hexagon, no HTPNollama.cpp's experimental Hexagon backend, compiled yourself — Snapdragon section above
Apple Silicon M-seriesThe GPU, via MetalNo — and no runtime anywhere doesNothing reaches the ANE: does Ollama use the Apple Neural Engine?

Three things people mistake for NPU acceleration in LM Studio

A GPU offload slider that visibly works. Pushing layers onto the GPU and watching tokens per second jump is real acceleration — from the integrated graphics. On a Ryzen AI or Lunar Lake laptop that is the Radeon or Arc block sharing the same memory pool, and it is the correct thing to be using. It is just not the NPU.

A runtime named after your CPU vendor. ROCm and Vulkan are graphics APIs. A build labelled for AMD or for Intel is labelled for that vendor's GPU, and the naming does a lot of accidental damage on a chip where the CPU, GPU and NPU all carry the same brand.

Task Manager showing NPU activity while LM Studio is open. Windows Studio Effects — background blur, eye contact, auto-framing — is an NPU workload, and it runs whenever your camera is on. That graph can be busy while your model is running entirely on the CPU two windows away. Close the camera app and watch the NPU line drop to zero while your tokens keep streaming; that is the test.

What to do if you want LM Studio's convenience and the NPU

Run both, and be clear about which silicon does what. flm serve exposes an OpenAI-compatible endpoint on port 52625, so any client that already points at a local base URL — Open WebUI, a VS Code extension, your own script — can talk to the NPU while LM Studio keeps handling the models you want on the GPU. That is two applications instead of one, which is genuinely worse ergonomics, and it is the honest state of things in August 2026: there is no single desktop app that gives you a model library, a chat window and NPU execution in the same process.

If none of that appeals, the answer is not a workaround, it is a different expectation. Use LM Studio on the iGPU, which works well, and read the measure-it-yourself procedure before assuming the NPU would have been faster anyway.


Measure It Yourself

The NPU vs iGPU vs CPU comparison on one machine is the number everybody wants, and we are not going to fabricate it.

We do not own a Kraken Point laptop, a Lunar Lake ultrabook and a Snapdragon X device, and a three-way tokens-per-second table assembled from three different vendors' marketing pages would be worse than no table at all — different chips, different quants, different context lengths, different power modes. So here is the ten-minute procedure to produce it for your machine, which is the only one that matters:

  1. Fix the variables first. Same model, same parameter count, same quantisation family, same prompt, same context length, plugged in, same power profile. Changing two things at once is how these comparisons become meaningless.
  2. NPU number. On AMD: flm run llama3.2:3b, then /verbose in the session to get performance reporting. On Intel: time the OpenVINO GenAI pipeline with device "NPU". On Snapdragon: llama-bench with the Hexagon backend enabled.
  3. iGPU number. Same model through Lemonade on the iGPU backend, or Ollama/llama.cpp with GPU offload, or on Intel the same OpenVINO pipeline with device "GPU".
  4. CPU number. Same model, offload disabled.
  5. Record watts and battery, not just tok/s. Unplug and run a fixed workload on each backend, and note the battery percentage drop. This is the axis on which NPUs are designed to win, and the axis every TOPS marketing table ignores.
  6. Report prefill and decode separately. FastFlowLM's own published split — high prefill, moderate decode — suggests the two numbers behave very differently on NPUs. A single averaged tok/s hides that.

If you run this on your own hardware, we would genuinely like to see it. The absence of this table on the open web as of August 2026 is the real gap, not a shortage of TOPS charts.


What NPUs Are Actually Good At

Buy the laptop for its memory and its GPU; treat the NPU as a bonus that pays off in battery life and background work.

Three things follow from everything above:

The NPU frees the GPU. If a local model is running as an always-on assistant, dictation engine or autocomplete, having it on the NPU means the iGPU stays available for the display, video calls and anything else. That is a real quality-of-life win that a raw tok/s comparison will never show.

Memory still decides what you can run. No NPU changes the fact that an 8B model at 4-bit needs roughly 5GB of memory to hold weights, plus KV cache. On a shared-memory laptop that comes out of the same pool as everything else. Model size limits on these machines are a memory story, and our laptop GPU VRAM guide and honest laptop guide cover the arithmetic.

TOPS is a purchasing spec, not a performance prediction. It rates peak arithmetic on the NPU block. Whether any runtime can reach it is a software question, and on two of these three platforms the answer as of August 2026 is "partially" and "if you compile it yourself". If you are choosing hardware rather than fixing hardware you own, the discrete-GPU comparison in Copilot+ PC vs RTX for local AI is the more useful read.


Verdict

  1. Ollama will never light up your NPU. No NPU backend exists in it. Stop looking for the setting.
  2. Neither will LM Studio, and its runtime picker is why you thought otherwise. Every runtime in that list is a CPU or GPU engine; its requirements page asks for AVX2 and VRAM and never mentions an NPU. Swapping runtimes moves work between the processor and the graphics block.
  3. AMD Ryzen AI owners: install FastFlowLM. XDNA2 chips only (Strix, Strix Halo, Kraken, Gorgon Point), NPU driver 32.0.203.304+, MSI installer, Ollama-style commands, now shipping under ROCm. This is the one path that feels finished.
  4. Intel Core Ultra owners: use OpenVINO GenAI, not ipex-llm. ipex-llm is archived with a no-support banner and security flags. Expect an optimum-cli export step, INT4 symmetric weights, and a default 1024-token prompt ceiling from the static-shape design.
  5. Snapdragon X owners: it works, but you are compiling it. llama.cpp's Hexagon backend is real, experimental, Q4_0-tested, and needs the Hexagon SDK. The CPU path is the sane default until that changes.
  6. Mac owners: stop looking entirely. The ANE is Core ML-only with no public low-level API, and MLX declined to support it on those grounds. The full explanation is here; Metal is your accelerator and it is a good one.
  7. Do not trust any NPU-vs-iGPU claim without a method attached — including ours. We have not measured all three, we said so, and we gave you the procedure instead.

The honest summary: NPU LLM inference on laptops went from "vapourware" to "one platform is genuinely usable" in about eighteen months. If you own a Ryzen AI machine, that is good news you can act on today. If you own the other two, the ceiling is still software, not silicon.


Sources


FAQ

🎯
AI Learning Path

Got the hardware sorted? Now build on it.

You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Once your hardware is sorted

Decide before you spend a thousand pounds

The AI Hardware course sizes your build properly — VRAM ladder, real bottlenecks, budget builds — and Pick the Right Model tells you what to run on it.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsNPURyzen AILunar LakeSnapdragon XFastFlowLMOpenVINOllama.cpp

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on Local AI Hardware
See the full AI Hardware Guide 2026 guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

Why is my NPU at 0% when I run Ollama?

Because Ollama has no NPU backend. Its hardware documentation lists NVIDIA GPUs, AMD Radeon via ROCm, Apple Metal and Vulkan — no AMD XDNA, no Intel Core Ultra NPU, no Qualcomm Hexagon. Nothing you configure in Ollama will change this. The NPU is a separate accelerator with its own driver and its own compiled kernels; reaching it requires software written specifically for it, which today means FastFlowLM or Lemonade on AMD, OpenVINO GenAI on Intel, and the experimental Hexagon backend in llama.cpp on Snapdragon.

Can I run an LLM on an AMD Ryzen AI NPU?

Yes, and it is the easiest of the three. FastFlowLM supports all Ryzen AI chips with XDNA2 NPUs — Strix, Strix Halo, Kraken and Gorgon Point — with a Windows MSI installer and Linux support added in March 2026. The command set deliberately mirrors Ollama: flm run llama3.2:1b, flm list, flm serve for an OpenAI-compatible server on port 52625. You need NPU driver 32.0.203.304 or newer. As of v1.0.0 the project ships under AMD ROCm, which is a meaningful signal about its future.

Is the Intel NPU usable for LLMs on Lunar Lake?

Yes, through OpenVINO GenAI, but not through the tool most guides point at. Intel archived ipex-llm on January 28, 2026 with a banner stating Intel will not provide or guarantee maintenance, bug fixes, new releases or updates, and flagging known security issues — so treat every ipex-llm NPU tutorial as dead. The maintained route is OpenVINO GenAI: export the model with optimum-cli to INT4 symmetric channel-wise weights, install NPU driver 32.0.100.3104 or newer, and run the GenAI LLM pipeline on device NPU.

Does the Snapdragon X Elite NPU run LLMs?

Experimentally. llama.cpp has a Hexagon NPU backend documented under docs/backend/snapdragon, covering Android arm64 and native Windows 11 arm64, with Hexagon architecture versions v73 through v81. It is labelled experimental, the demonstrated quantisation is Q4_0, and there is no installer — you build it yourself with the Hexagon SDK, and on Windows also Visual Studio, LLVM/Clang, CMake and the OpenCL SDK. The docs report roughly 51.5 tokens/second token generation for Llama-3.2-1B in their tg64 benchmark. If you want a click-and-go experience on a Snapdragon laptop today, run on the CPU instead.

Does LM Studio support the Intel or Ryzen AI NPU?

No. LM Studio lets you pick a runtime, which is why people assume one of the options is the NPU, but every entry in that list is a CPU or GPU engine. Its own system requirements page asks for AVX2 on x64 and around 4GB of dedicated VRAM, mentions Apple Silicon and "ARM (Snapdragon X Elite) based systems", and never mentions an NPU on any platform. Switching runtimes moves work between the CPU and the graphics block. On a Ryzen AI laptop the Radeon graphics and the XDNA NPU are separate devices behind separate drivers, so a ROCm or Vulkan runtime cannot reach the NPU by definition. For the NPU you leave LM Studio: FastFlowLM or Lemonade on AMD, OpenVINO GenAI on Intel.

Does the Apple Neural Engine run LLMs?

No, and the reason is different from the Windows story. The ANE is reachable only through Core ML — there is no public low-level API to compile and dispatch a kernel the way you can with Metal or CUDA. MLX maintainers closed the ANE support request in December 2023 saying they had no plans to support it while it remains a closed-source API. Ollama, llama.cpp and MLX all target the GPU on Apple Silicon. Our separate page on Ollama and the Apple Neural Engine covers the whole chain and gives you the powermetrics command to see what your own Mac is really doing.

Is the NPU faster than the integrated GPU for LLMs?

Nobody publishes an honest three-way NPU vs iGPU vs CPU comparison on the same machine, and we do not own all three laptops, so we will not invent one. What we can tell you is how to produce it yourself in about ten minutes: run FastFlowLM with /verbose for the NPU number, then llama-bench or Ollama with the same GGUF quant on the iGPU and CPU, same prompt, same context length. Report all three. The architectural pitch for NPUs is sustained throughput at low power rather than peak speed, so measure watts and battery drain alongside tokens per second — that is where the answer usually lives.

What does the "50 TOPS" number on my Copilot+ laptop actually buy me?

TOPS is a peak arithmetic rating for the NPU block, and it says nothing about whether any LLM software can reach it. That is the whole gap this page exists to close: a 50-TOPS NPU with no runtime targeting it produces exactly zero tokens per second. TOPS became a purchasing spec because Microsoft set a 40-TOPS floor for the Copilot+ badge, not because it predicts local LLM performance. Judge a laptop for local AI on memory bandwidth, memory capacity and which runtimes support the silicon.

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks for local models — CPU and GPU paths that just work. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: October 4, 2026🔄 Last Updated: October 4, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Got the hardware sorted? Now build on it.

You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators