Run an LLM on Your Laptop\'s NPU: What Actually Works
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Got the hardware sorted? Now build on it. You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.
Short answer: your NPU shows 0% because Ollama has no NPU backend at all — its own hardware docs list NVIDIA, AMD ROCm, Apple Metal and Vulkan, and nothing else. LM Studio does not have one either; its runtime picker only chooses between CPU and GPU engines. To actually reach the NPU you need software written for that specific silicon: FastFlowLM or Lemonade on AMD Ryzen AI (works today, closest thing to an Ollama experience), OpenVINO GenAI on Intel Core Ultra (works, but you must export the model first), and the experimental Hexagon backend in llama.cpp on Snapdragon X (build it yourself, no installer). Intel's friendlier ipex-llm route is archived and should not be used.
That is the whole answer. The rest of this page is the per-chip detail, the exact commands, the driver versions, a section on why swapping runtimes in LM Studio changes nothing, and one section on the comparison everybody wants and nobody has published.
Why Your NPU Sits at 0%
It is not a setting. Ollama, every runtime LM Studio ships and llama.cpp's mainline builds do not target NPUs, so there is nothing to enable. LM Studio is the confusing one because it does hand you a runtime chooser — that gets its own section below.
An NPU is not a small GPU. It has its own driver stack, its own memory-movement model, and it runs precompiled kernels rather than the general-purpose shader or CUDA code that GPU inference engines emit. That means a runtime cannot "also support NPU" as a checkbox — someone has to write and ship compiled kernels per operator, per quantisation format, per NPU generation. That work has happened for a handful of projects and not for the ones you already have installed.
Ollama's hardware documentation is explicit about the accelerators it supports: NVIDIA GPUs at compute capability 5.0+, AMD Radeon via ROCm, Apple Metal, and additional GPUs via Vulkan on Windows and Linux. AMD Ryzen AI parts appear in that documentation as Radeon integrated graphics, which is the iGPU — not the XDNA NPU. Intel Core Ultra NPU and Qualcomm Hexagon are absent entirely.
So when you bought a machine advertised at 50 TOPS and Task Manager's NPU graph stays flat while the CPU cooks, nothing is broken. You are simply running software that was never pointed at that block of silicon. If you want the buying-side view of what those TOPS ratings mean across vendors, that is our NPU comparison; this page is about making the thing move.
Mac owners, take this exit now. The Apple Neural Engine is the same story with a harder ceiling: it is reachable only through Core ML, there is no public low-level API to compile and dispatch kernels the way you can with Metal or CUDA, and MLX's maintainers closed the ANE request in December 2023 saying they have no plans while it stays a closed-source API. Nothing on this page will light up your ANE, because nothing does. The full chain — Ollama, llama.cpp, MLX, and the powermetrics command to watch your own Mac — is in does Ollama use the Apple Neural Engine?. Everything from here down is about Windows laptops.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
Verdict Per Chip
AMD is usable today, Intel is usable with an extra export step, Snapdragon is a build-it-yourself project.
| Your chip | Can you run an LLM on the NPU? | Software | Friction |
|---|---|---|---|
| AMD Ryzen AI (XDNA2) — Strix, Strix Halo, Kraken, Gorgon Point | Yes | FastFlowLM, or Lemonade (which uses FastFlowLM as its NPU backend) | Low — Windows MSI installer, Ollama-style CLI |
| Intel Core Ultra Series 2 (Lunar Lake) and Core Ultra NPUs | Yes, with caveats | OpenVINO GenAI | Medium — model must be exported to INT4 OpenVINO IR first; static-shape limits |
| Qualcomm Snapdragon X | Experimental | llama.cpp Hexagon backend | High — compile from source with the Hexagon SDK |
| Ryzen AI 300 with older XDNA1 (Phoenix/Hawk Point) | Not via FastFlowLM | — | FastFlowLM's stated support is XDNA2 parts |
| Apple Silicon M-series (Neural Engine) | No — and not through any runtime | — | Blocked at the API layer, not the silicon: ANE vs Metal, explained |
One cross-cutting note before the details: on all three platforms the iGPU and CPU paths still work normally and are far less trouble. Running on the NPU is a deliberate choice you make for power or for freeing the GPU, not a prerequisite for local AI on these laptops. If you just want models running tonight, run them on the CPU and come back to this page later.
AMD Ryzen AI (XDNA2)
This is the one that works like you hoped. FastFlowLM installs from an MSI and its CLI is a deliberate Ollama clone.
FastFlowLM (1.8k stars) describes itself as Ollama built for AMD Ryzen AI NPUs. Supported silicon, per its README: all Ryzen AI series chips with XDNA2 NPUs — Strix, Strix Halo, Kraken and Gorgon Point. Windows gets an flm-setup.msi installer; Linux support arrived in March 2026. Version 1.0.0 landed in August 2026 under the headline "First Release Under ROCm" — the project now ships as part of AMD's ROCm organisation, which matters more for its longevity than any benchmark on this page. Licensing: MIT for the orchestration code and CLI, with the NPU kernels free for any use including commercial.
Requirement to check first: NPU driver 32.0.203.304 or newer (the project recommends the .311 build). Older Ryzen AI driver packages will not work. You also need internet access on first run so it can pull the compiled NPU kernels for your model.
flm list # models available
flm run llama3.2:1b # interactive, on the NPU
flm serve llama3.2:1b # OpenAI-compatible server on port 52625
Inside a session, /verbose turns on performance reporting and /bye exits — again, straight out of the Ollama playbook. Because flm serve speaks the OpenAI API on port 52625, anything you already point at a local endpoint (Open WebUI, a VS Code extension, your own script) works by changing the base URL.
Model coverage is broad for a project this young: the documented families include LLaMA, Qwen, Gemma, gpt-oss, Phi, DeepSeek, LiquidAI LFM, Whisper and EmbeddingGemma, with the docs claiming full context length support across the list. If gpt-oss is what you are after, our gpt-oss model page covers the weights themselves.
The published speed numbers
FastFlowLM publishes per-model benchmark pages. Their LLaMA results, measured on an AMD Ryzen AI 7 350 (Kraken Point) with 32GB DRAM, FastFlowLM v0.9.30, default performance mode:
| Model | Decode (tok/s @ 1k context) | Prefill (tok/s @ 1k prompt) |
|---|---|---|
| LLaMA 3.2 1B | 64.5 | 1,686 |
| LLaMA 3.2 3B | 26.3 | 766 |
| LLaMA 3.1 8B | 12.8 | 403 |
Source: FastFlowLM benchmark documentation. These are vendor-published figures on one specific chip, with no iGPU or CPU baseline alongside them — read the measure-it-yourself section before drawing conclusions.
Take the shape of that table seriously even if you discount the absolute values: prefill is where the NPU looks strong (1,686 tok/s ingesting a prompt on a 1B model is a lot for a thin-and-light), while decode at 12.8 tok/s for an 8B is readable-but-not-fast. That profile — fast at chewing through context, moderate at generating — is worth knowing when you decide what to use it for.
Lemonade: the same NPU, a wider net
Lemonade (5.4k stars, Apache-2.0) is the broader runtime: it serves LLMs across NPUs, GPUs and CPUs — XDNA2 NPU, NVIDIA GPUs from Turing on, AMD RDNA3/4 including the Strix Halo iGPU, Apple Silicon, and x86_64/ARM64 CPUs. It also handles speech-to-text, TTS and image generation, not just text. Its NPU support is FastFlowLM under the hood: the v11.0.0 release notes describe the FastFlowLM NPU backend auto-installing on Linux rather than needing a manual system package. v11.6.0 shipped 14 August 2026, so it is moving fast.
lemonade backends # what your machine can actually use
lemonade run Gemma-4-E2B-it-GGUF
Pick Lemonade if you want one runtime that falls back gracefully across NPU, iGPU and CPU, or you want the non-text modalities. Pick FastFlowLM directly if you only care about the NPU and want the smallest possible stack. For the Strix Halo class of machine specifically, see our AI Max+ 395 guide and the best Strix Halo mini PC roundup — on those boxes the 128GB unified memory pool is a bigger story than the NPU.
Intel Core Ultra / Lunar Lake
The NPU works, but the tutorial you found is probably for a dead project. Intel archived ipex-llm on 28 January 2026.
Start with the bad news, because it will save you an evening. The repository intel/ipex-llm — 8.9k stars, the project behind nearly every "run an LLM on your Intel NPU" blog post from 2024 and 2025 — is archived. Its banner reads:
Intel will not provide or guarantee development of or support for this project, including but not limited to, maintenance, bug fixes, new releases or updates.
GitHub additionally flags it as having known security issues. Its NPU quickstarts are still visible and still indexed, which is exactly why people keep landing on them. Do not build anything new on it.
The maintained path is OpenVINO GenAI. It is more work than ollama run, but it is documented, current and supported.
# 1. Export and compress the model to OpenVINO IR, INT4 symmetric channel-wise
optimum-cli export openvino -m meta-llama/Meta-Llama-3.1-8B-Instruct \
--weight-format int4 --sym --ratio 1.0 --group-size -1 \
Meta-Llama-3.1-8B-Instruct
Then load that exported directory with an OpenVINO GenAI LLM pipeline and pass "NPU" as the device. Things the documentation is clear about, and that you should plan around:
- NPU driver: install the latest; the troubleshooting guidance names 32.0.100.3104 or newer when execution fails.
- Weight formats accepted on NPU: INT4 symmetric channel-wise, INT4 group quantisation at group size 128, NF4 (Series 2 NPUs and later only), and INT8 for Whisper models. You cannot just point it at a random GGUF — this is a different model format entirely.
- Static shapes. The NPU pipeline compiles for fixed shapes, which is what makes it fast and also what constrains it: the documented defaults are a maximum input prompt of 1024 tokens and a minimum response length of 128 tokens.
- RAM. Intel documents that Core Ultra Series 2 systems may need more than 16GB of RAM to process prompts longer than 1024 tokens with models above 7B parameters. On a 16GB Lunar Lake ultrabook, plan on smaller models.
That 1024-token default prompt ceiling is the honest headline for Intel. It is adjustable, but the static-shape design means long-context chat is not what this path is built for. Summarising a short email, running a local autocomplete, powering a small assistant: fine. Feeding it a 30-page PDF: use the iGPU. Intel's discrete cards are a different and much less constrained story — see Intel Arc B580 for local AI.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
Snapdragon X
There is a real Hexagon NPU backend in llama.cpp, it is labelled experimental, and there is no installer.
llama.cpp carries Snapdragon documentation under docs/backend/snapdragon covering three backends on these devices: CPU, Adreno GPU via OpenCL, and Hexagon NPU (HTP). Facts worth having before you commit a weekend:
- Platforms: Android arm64, and native Windows 11 arm64 builds. Linux appears in the Docker toolchain instructions.
- Hexagon versions: the docs reference v73, v75, v79 and v81, varying by hardware generation.
- Quantisation: the worked examples use Q4_0 (e.g.
Llama-3.2-1B-Instruct-Q4_0.gguf). Treat Q4_0 as the tested format. - Memory model: models under 4B fit in one Hexagon session; an 8B needs two; a 20B needs four. You control this with
GGML_HEXAGON_NDEV. - How offload works: the docs note the Hexagon NPU "behaves as a GPU device when it comes to
-ngland other offload-related options", so the flags you already know carry over. - Debugging:
GGML_HEXAGON_PROFILEandGGML_HEXAGON_VERBOSEenvironment variables. - Published speed: roughly 51.5 tok/s token generation for Llama-3.2-1B in their tg64 benchmark. That is one model at one size on unstated silicon — take it as evidence the path works, not as a spec.
The cost is the build. The Docker toolchain image bundles the Android NDK, OpenCL SDK, Hexagon SDK and CMake; a native Windows-on-Snapdragon build wants Visual Studio (for the Windows SDK and Driver Kit), LLVM/Clang, CMake, Hexagon SDK 6.6.0.0 and the Adreno OpenCL SDK 2.3.2. There is no MSI and no brew install.
Our honest recommendation for Snapdragon X owners: unless you enjoy toolchains, run models on the CPU — these are competent ARM cores and llama.cpp's CPU path on Windows-on-ARM is mature — and revisit the Hexagon backend when it loses the experimental label. Nothing about it is fake; it is just not a consumer install yet.
Does LM Studio Support the Intel or Ryzen AI NPU?
LM Studio has no NPU runtime on any platform. The list you are choosing from contains CPU engines and GPU engines, so switching runtimes moves the model between the processor and the graphics block — never onto the Intel, Ryzen AI or Snapdragon NPU.
This deserves its own answer because LM Studio is the one app that makes the mistake reasonable. Ollama hides its backend selection entirely, so nobody expects a knob. LM Studio puts a runtime chooser right in front of you, and a screen full of selectable engines invites the obvious guess: one of these must be the NPU. None of them is.
The app's own requirements are the tell. LM Studio's system requirements ask for "AVX2 instruction set support is required (for x64)" and "at least 4GB of dedicated VRAM is recommended" — one CPU requirement and one GPU requirement. Apple Silicon gets a chip line. Windows on ARM is listed as "ARM (Snapdragon X Elite) based systems". An NPU is not mentioned on any of the three platforms, and neither is a TOPS floor. A vendor that shipped NPU acceleration would say so on that page.
What each runtime actually targets on your machine
| Your machine | What the runtime list is offering you | NPU reached? | Where the NPU path actually lives |
|---|---|---|---|
| Intel Core Ultra (Lunar Lake, Arrow Lake) | CPU builds, plus the Arc integrated graphics as a GPU device | No | OpenVINO GenAI, a separate toolchain with its own export step — Intel section above |
| AMD Ryzen AI 300 / AI Max (XDNA2) | CPU builds, plus the Radeon integrated graphics as a GPU device | No. Radeon graphics and the XDNA NPU are two different devices behind two different drivers — a graphics API cannot address the NPU | FastFlowLM, or Lemonade — AMD section above |
| Qualcomm Snapdragon X Elite | The ARM64 CPU. The requirements page names the architecture and stops — no Hexagon, no HTP | No | llama.cpp's experimental Hexagon backend, compiled yourself — Snapdragon section above |
| Apple Silicon M-series | The GPU, via Metal | No — and no runtime anywhere does | Nothing reaches the ANE: does Ollama use the Apple Neural Engine? |
Three things people mistake for NPU acceleration in LM Studio
A GPU offload slider that visibly works. Pushing layers onto the GPU and watching tokens per second jump is real acceleration — from the integrated graphics. On a Ryzen AI or Lunar Lake laptop that is the Radeon or Arc block sharing the same memory pool, and it is the correct thing to be using. It is just not the NPU.
A runtime named after your CPU vendor. ROCm and Vulkan are graphics APIs. A build labelled for AMD or for Intel is labelled for that vendor's GPU, and the naming does a lot of accidental damage on a chip where the CPU, GPU and NPU all carry the same brand.
Task Manager showing NPU activity while LM Studio is open. Windows Studio Effects — background blur, eye contact, auto-framing — is an NPU workload, and it runs whenever your camera is on. That graph can be busy while your model is running entirely on the CPU two windows away. Close the camera app and watch the NPU line drop to zero while your tokens keep streaming; that is the test.
What to do if you want LM Studio's convenience and the NPU
Run both, and be clear about which silicon does what. flm serve exposes an OpenAI-compatible endpoint on port 52625, so any client that already points at a local base URL — Open WebUI, a VS Code extension, your own script — can talk to the NPU while LM Studio keeps handling the models you want on the GPU. That is two applications instead of one, which is genuinely worse ergonomics, and it is the honest state of things in August 2026: there is no single desktop app that gives you a model library, a chat window and NPU execution in the same process.
If none of that appeals, the answer is not a workaround, it is a different expectation. Use LM Studio on the iGPU, which works well, and read the measure-it-yourself procedure before assuming the NPU would have been faster anyway.
Measure It Yourself
The NPU vs iGPU vs CPU comparison on one machine is the number everybody wants, and we are not going to fabricate it.
We do not own a Kraken Point laptop, a Lunar Lake ultrabook and a Snapdragon X device, and a three-way tokens-per-second table assembled from three different vendors' marketing pages would be worse than no table at all — different chips, different quants, different context lengths, different power modes. So here is the ten-minute procedure to produce it for your machine, which is the only one that matters:
- Fix the variables first. Same model, same parameter count, same quantisation family, same prompt, same context length, plugged in, same power profile. Changing two things at once is how these comparisons become meaningless.
- NPU number. On AMD:
flm run llama3.2:3b, then/verbosein the session to get performance reporting. On Intel: time the OpenVINO GenAI pipeline with device"NPU". On Snapdragon:llama-benchwith the Hexagon backend enabled. - iGPU number. Same model through Lemonade on the iGPU backend, or Ollama/llama.cpp with GPU offload, or on Intel the same OpenVINO pipeline with device
"GPU". - CPU number. Same model, offload disabled.
- Record watts and battery, not just tok/s. Unplug and run a fixed workload on each backend, and note the battery percentage drop. This is the axis on which NPUs are designed to win, and the axis every TOPS marketing table ignores.
- Report prefill and decode separately. FastFlowLM's own published split — high prefill, moderate decode — suggests the two numbers behave very differently on NPUs. A single averaged tok/s hides that.
If you run this on your own hardware, we would genuinely like to see it. The absence of this table on the open web as of August 2026 is the real gap, not a shortage of TOPS charts.
What NPUs Are Actually Good At
Buy the laptop for its memory and its GPU; treat the NPU as a bonus that pays off in battery life and background work.
Three things follow from everything above:
The NPU frees the GPU. If a local model is running as an always-on assistant, dictation engine or autocomplete, having it on the NPU means the iGPU stays available for the display, video calls and anything else. That is a real quality-of-life win that a raw tok/s comparison will never show.
Memory still decides what you can run. No NPU changes the fact that an 8B model at 4-bit needs roughly 5GB of memory to hold weights, plus KV cache. On a shared-memory laptop that comes out of the same pool as everything else. Model size limits on these machines are a memory story, and our laptop GPU VRAM guide and honest laptop guide cover the arithmetic.
TOPS is a purchasing spec, not a performance prediction. It rates peak arithmetic on the NPU block. Whether any runtime can reach it is a software question, and on two of these three platforms the answer as of August 2026 is "partially" and "if you compile it yourself". If you are choosing hardware rather than fixing hardware you own, the discrete-GPU comparison in Copilot+ PC vs RTX for local AI is the more useful read.
Verdict
- Ollama will never light up your NPU. No NPU backend exists in it. Stop looking for the setting.
- Neither will LM Studio, and its runtime picker is why you thought otherwise. Every runtime in that list is a CPU or GPU engine; its requirements page asks for AVX2 and VRAM and never mentions an NPU. Swapping runtimes moves work between the processor and the graphics block.
- AMD Ryzen AI owners: install FastFlowLM. XDNA2 chips only (Strix, Strix Halo, Kraken, Gorgon Point), NPU driver 32.0.203.304+, MSI installer, Ollama-style commands, now shipping under ROCm. This is the one path that feels finished.
- Intel Core Ultra owners: use OpenVINO GenAI, not ipex-llm. ipex-llm is archived with a no-support banner and security flags. Expect an
optimum-cliexport step, INT4 symmetric weights, and a default 1024-token prompt ceiling from the static-shape design. - Snapdragon X owners: it works, but you are compiling it. llama.cpp's Hexagon backend is real, experimental, Q4_0-tested, and needs the Hexagon SDK. The CPU path is the sane default until that changes.
- Mac owners: stop looking entirely. The ANE is Core ML-only with no public low-level API, and MLX declined to support it on those grounds. The full explanation is here; Metal is your accelerator and it is a good one.
- Do not trust any NPU-vs-iGPU claim without a method attached — including ours. We have not measured all three, we said so, and we gave you the procedure instead.
The honest summary: NPU LLM inference on laptops went from "vapourware" to "one platform is genuinely usable" in about eighteen months. If you own a Ryzen AI machine, that is good news you can act on today. If you own the other two, the ceiling is still software, not silicon.
Sources
- Ollama GPU documentation — the supported-accelerator list, with no NPU entry
- FastFlowLM — XDNA2 support list, driver requirement, CLI, licence, ROCm release note
- FastFlowLM benchmarks — Ryzen AI 7 350 decode/prefill figures
- lemonade-sdk/lemonade — backend coverage, NPU-via-FastFlowLM, release cadence
- intel/ipex-llm — archive banner and support statement
- OpenVINO GenAI on NPU documentation — driver version, INT4 weight formats, export command, static-shape limits
- llama.cpp Snapdragon backend docs — Hexagon versions, Q4_0, session model, environment variables, tg64 figure
- LM Studio system requirements — the AVX2 and dedicated-VRAM requirements, the Snapdragon X Elite ARM line, and the complete absence of any NPU mention
FAQ
Got the hardware sorted? Now build on it.
You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.
Decide before you spend a thousand pounds
The AI Hardware course sizes your build properly — VRAM ladder, real bottlenecks, budget builds — and Pick the Right Model tells you what to run on it.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARLocal AI Hardware Requirements (2026): Complete Guide
- AI Hardware Requirements: CPU, GPU and RAM for Beginners
- AI RAM Requirements 2026: How Much for 7B, 13B, 70B Models?
- AI Server Build Under $1,500: Parts List and What Fits
- AMD GPU Not Supported by ROCm? HSA_OVERRIDE Values
- AMD MI50 32GB for Local LLMs: The Used VRAM King, Honestly
- AMD Ryzen AI Max+ 395 (Strix Halo) for Local AI 2026
- Apple M4 for Local AI: Mac Studio + MacBook Guide (2026)
- Benchmark Your Local AI Setup: tok/s, TTFT, VRAM
- Best GPU for AI Video Generation: By VRAM Tier (2026)
Comments (0)
No comments yet. Be the first to share your thoughts!