★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Troubleshooting

Ollama Not Using GPU: 8 Causes, One Test for Each

September 13, 2026
17 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Ollama’s running. Here’s what to build with it. Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.

Start free
Or own it for life — Lifetime $149, pay once

Short answer: run ollama ps while the model is loaded and read the PROCESSOR column. "100% GPU" means it works. "100% CPU" means Ollama never found a usable GPU — check the server log for no compatible GPUs were discovered, then go to causes 2-8. Anything in between, like 48%/52% CPU/GPU, means the GPU was found and the model did not fit — that is cause 1, and it has nothing to do with your drivers. Those are two completely different problems and almost every guide on the internet mashes them together into "check your drivers."

Written against Ollama v0.32.14 (released 2026-08-15), the current release at the time of writing, with every log line and environment variable below traced to Ollama's own docs or to a linked issue in its tracker. We call the version out because GPU discovery and its log output have been reworked more than once across releases — if you are on something older, the exact log lines may differ.


Start Here: The One Command That Splits the Problem

ollama ps is the diagnostic. Not nvidia-smi.

Load a model, leave it loaded, and run this in a second terminal:

# terminal 1 — keep this open so the model stays resident
ollama run qwen3:8b "write a haiku about VRAM"

# terminal 2 — while the model is still loaded
ollama ps

You get something like:

NAME        ID              SIZE      PROCESSOR          UNTIL
qwen3:8b    <model-id>      6.5 GB    48%/52% CPU/GPU    4 minutes from now

Ollama's own FAQ documents the three states of that PROCESSOR column: 100% GPU (entirely on the GPU), 100% CPU (entirely in system memory), and a mixed percentage meaning the model "was loaded partially onto both the GPU and into system memory."

Why not nvidia-smi? Because it answers a different question. It tells you a process exists and how much VRAM is allocated — it does not tell you how many layers are on the card, and utilisation sampled between token batches can read near zero on a perfectly healthy run. People stare at 0% utilisation and conclude the GPU is unused when the real problem is that only half the layers are there.

Note the SIZE column too. That number is weights plus KV cache plus compute buffers, not the download size of the model. It is routinely 1-4 GB larger than the file you pulled, and that gap is the single most common reason a "4 GB model" does not fit on a 12 GB card.

Now branch:

ollama ps saysYour problem isGo to
100% GPUNot a GPU-detection problem. If it is still slow, it is model size, quantisation or thermalsSlow local LLM fix
A split, e.g. 48%/52%The GPU works. The model did not fitCause 1
100% CPUOllama found no usable GPURead the log line below, then causes 2-8

Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

"No Compatible GPUs Were Discovered" — What That Log Line Actually Means

no compatible GPUs were discovered means discovery ran at server start and came back with an empty list. Ollama never attempted to use a card, so nothing you change at run time will help — and it is logged at INFO, not ERROR, which is why people grep their logs and find nothing.

The line looks like this, and the exact wording is stable enough that it is worth searching for verbatim:

time=<timestamp> level=INFO source=gpu.go:<line> msg="looking for compatible GPUs"
time=<timestamp> level=INFO source=gpu.go:<line> msg="no compatible GPUs were discovered"

A GitHub search of the ollama/ollama tracker on 2026-08-22 returns 152 issues and pull requests containing that exact string, against 592 for the looking for compatible GPUs line that precedes it. The pair is the single most-reported symptom of this whole problem, and the reason it gets misdiagnosed is the log level: grep -i error and grep -i fail both miss it completely.

# the grep that actually catches it, on all three platforms
journalctl -u ollama --no-pager | grep -E "looking for compatible GPUs|no compatible GPUs|amdgpu is not supported|compute capability"

# macOS
grep -E "compatible GPUs|metal|amdgpu" ~/.ollama/logs/server.log

# Docker
docker logs <container-name> 2>&1 | grep -E "compatible GPUs|amdgpu is not supported"

The line immediately above it is the one that names your cause. Discovery logs why each backend was rejected before it prints the empty result. These are the three you are most likely to see, quoted as they appear in the tracker:

Preceding lineWhat it meansGo to
level=WARN source=amd_linux.go:443 msg="amdgpu detected, but no compatible rocm library found. Either install rocm v6, or follow manual install instructions..."The kernel driver is loaded but the ROCm userspace libraries are missing or the wrong major versionCause 2, then cause 8
level=WARN source=amd_linux.go:378 msg="amdgpu is not supported (supported types:[gfx1030 gfx1100 gfx1101 gfx1102 gfx900 gfx906 gfx908 gfx90a gfx940 gfx941 gfx942])" gpu_type=gfx1031 gpu=0ROCm is installed and working. Your chip's gfx target is simply not in the list of code objects Ollama shippedCause 8
Nothing at all between looking for compatible GPUs and the empty resultThe process could not see a device node in the first place — permissions, container, or a visible-devices filterCauses 5 and 7

Note the second line prints two things you need: gpu_type= is your card's target, and the bracketed list is what Ollama actually built for. On Windows the same warning comes from amd_windows.go instead of amd_linux.go. Older Ollama builds format it as msg="amdgpu is not supported" gpu=0 gpu_type=gfx1035 library=/opt/rocm/lib supported_types="[...]" — same information, different punctuation.


Cause 1: Partial Offload — The KV Cache Ate Your VRAM

If PROCESSOR shows a split, this is your cause, and the fix is context length — not drivers. At Ollama's default 4096-token context a Qwen3-8B-class model needs about 0.56 GB of f16 KV cache. At 32768 tokens the same model needs about 4.5 GB. That is the difference between fitting on a 12 GB card and spilling layers to the CPU.

The KV cache is allocated per loaded model and scales linearly with context. The formula is not a mystery:

KV bytes per token = 2 (K and V) x layers x kv_heads x head_dim x bytes_per_element

Plug in published model configs and you get concrete numbers. Layer counts and KV-head counts below are from each model's own config.json on Hugging Face; the arithmetic is ours, at f16 (2 bytes):

ModelLayersKV headshead_dimKV @ 4096 ctxKV @ 32768 ctx
Qwen3-8B3681280.56 GB4.5 GB
Qwen3-32B6481281.0 GB8.0 GB

Two things fall out of that table. First, an 8× context increase is an 8× KV increase — there is no clever compression happening by default. Second, the 32B model burns 8 GB of KV cache alone at 32K context, before a single weight is loaded. That is why people with 24 GB cards find a 32B Q4 model (~20 GB of weights) suddenly half on the CPU the moment they raise the context.

The one-command test: reload with the default context and see if the split goes away.

# stop the server, then start it with an explicit small context
OLLAMA_CONTEXT_LENGTH=4096 ollama serve

If ollama ps now reads 100% GPU, you have confirmed it: context was the problem, not hardware.

The three fixes, in the order we would try them:

  1. Lower the context to what you actually need. Most chat use never touches 8K. Ollama's documented default is 4096 tokens, settable per-server with OLLAMA_CONTEXT_LENGTH, per-session with /set parameter num_ctx 4096, or per-request with "num_ctx" in the API.
  2. Quantise the KV cache. Ollama documents OLLAMA_KV_CACHE_TYPE with three values: f16 (default), q8_0 (about half the memory of f16) and q4_0 (about a quarter). Setting q8_0 turns that 8 GB Qwen3-32B cache into roughly 4 GB. This is a quality trade, and Ollama's own documentation says so — it notes that a quantised KV cache reduces memory at some cost to precision, and that the impact varies by model. Test both values on your own prompts before adopting either.
  3. Drop a quantisation level on the weights. Q4_K_M instead of Q6_K buys back more room than any tuning flag.
# half the KV cache, one line
OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_CONTEXT_LENGTH=8192 ollama serve

Before you buy anything, run the numbers for the exact model and context you want on our VRAM calculator, and cross-check against the Ollama model RAM/VRAM table.

One more thing that catches people: if two models are resident at once, both are holding VRAM. Set OLLAMA_MAX_LOADED_MODELS=1, or OLLAMA_KEEP_ALIVE=0 to unload immediately after each request. Ollama keeps models in memory for five minutes by default.


Cause 2: Your Driver Is Below the Floor

Ollama's hardware docs state driver 550+ as the general requirement, and driver 570+ for GPUs with compute capability 5.0-6.2. Below the floor, Ollama falls back to CPU silently while nvidia-smi keeps working perfectly — which is precisely why this one wastes so many evenings.

The one-command test:

nvidia-smi --query-gpu=name,driver_version,compute_cap --format=csv

One line, three answers: which card, which driver, which compute capability. Compare the driver number to the floor. If your card is Maxwell or Pascal (compute capability 5.0-6.2 — GTX 900 and GTX 10-series), the number you need is 570, not 550.

The trap is long-term-support distro drivers. A stable Ubuntu box that has been happily running an older branch for two years will show a healthy nvidia-smi and give you 100% CPU in Ollama. Upgrading the driver is the whole fix.

# Debian/Ubuntu — check what your distro actually offers before installing
apt-cache search '^nvidia-driver-[0-9]' | sort

After any driver change, reboot. Not "restart Ollama" — reboot. And on Linux, if you updated the kernel recently, the NVIDIA modules may need rebuilding before they load at all; our Ollama troubleshooting guide covers the DKMS side of that.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

Cause 3: The Card Is Below the Compute-Capability Cutoff

Ollama supports NVIDIA GPUs with compute capability 5.0 or higher. Below 5.0 there is no fix and no flag — the card will never be used, and everything else in this guide is a waste of your time. Get the number from the compute_cap field in the command above.

Here is the supported table as published in Ollama's hardware docs, so you can find your card:

Compute capabilityCardsDriver needed
12.1GB10 (DGX Spark)550+
12.0GeForce RTX 50-series, RTX PRO Blackwell550+
9.0H200, H100550+
8.9RTX 40-series, L4, L40, RTX 6000550+
8.6RTX 30-series, A-series professional550+
8.0A100, A30550+
7.5GTX/RTX 20-series, T-series, Quadro RTX550+
7.0TITAN V, V100550+
6.1TITAN Xp, TITAN X, GTX 10-series570+
6.0Tesla P100, Quadro GP100570+
5.2GTX 900-series, M-series Quadro/Tesla570+
5.0GTX 750-series, older Quadro/Tesla570+

Source: Ollama's docs/gpu.mdx, read on 2026-08-18 against the v0.32.x tree.

Worth being precise here, because the forums are not: Pascal is not dropped. GTX 10-series cards at compute capability 6.1 are still in the supported table. What changed is that they sit in the 570+ driver band, so an old driver strands them on CPU in a way that reads exactly like "unsupported." If you are running a 10-series card, check the driver before you believe anyone telling you the card is dead. And check the docs for your Ollama version — this matrix has moved before and can move again.

If your card is genuinely below 5.0, the honest answer is that CPU inference on a small model is your path: a 3B at Q4 is usable on a modern CPU, a 70B is not. The Ollama system requirements page has the RAM-to-model-size mapping for CPU-only boxes.


Cause 4: GPU Discovery Failed at Server Start

If the driver and the card are both fine, the next thing to read is the server log from the moment the server started — not from when you ran the model. GPU discovery happens once, at startup, and the failure is recorded there with a numeric CUDA error code.

The one-command test — read the log:

# Linux (systemd)
journalctl -u ollama --no-pager | grep -i -E "gpu|cuda|rocm|vulkan|library" | tail -40

# macOS
grep -i -E "gpu|metal|library" ~/.ollama/logs/server.log | tail -40

# Docker
docker logs <container-name> 2>&1 | grep -i -E "gpu|cuda|library" | tail -40

On Windows the log locations are documented as %LOCALAPPDATA%\Ollama for server logs, %LOCALAPPDATA%\Programs\Ollama for binaries and %HOMEPATH%\.ollama for models. To get verbose output there, quit the app from the tray menu first, then:

$env:OLLAMA_DEBUG="1"
& "ollama app.exe"

The error codes, straight from Ollama's troubleshooting docs:

CodeMeaningWhat it usually is
3not initializedThe CUDA/UVM driver stack did not come up — most often after a kernel or driver update without a reboot
46device unavailableSomething else has the device, or it is in an exclusive-compute mode
100no deviceThe process genuinely cannot see a GPU — permissions, container, or a CUDA_VISIBLE_DEVICES filter
999unknownThe catch-all. Treat it as "read dmesg", not as a specific fault

The documented first move for codes 3 and 999 on Linux is the UVM module:

sudo nvidia-modprobe -u
# or, to force a clean reload
sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm

For more detail than the log gives you, Ollama documents CUDA_ERROR_LEVEL=50 for detailed CUDA logging, and sudo dmesg | grep -i nvidia for anything the kernel is complaining about (Xid errors, falling-off-the-bus messages). On AMD, the equivalents are AMD_LOG_LEVEL=3 for HIP/ROCm detail and OLLAMA_DEBUG=1 during discovery; the docs note that "failure during GPU discovery" or bootstrap timeouts on AMD usually mean an outdated driver, with ROCm v7 as the current Linux requirement. If discovery instead prints amdgpu is not supported with a gpu_type= field, the driver is fine and you are in cause 8 — a gfx target Ollama did not build for.

# maximum-detail restart, one line
OLLAMA_DEBUG=1 CUDA_ERROR_LEVEL=50 ollama serve 2>&1 | tee /tmp/ollama-gpu.log

Cause 5: The Container or Service Cannot See the Device

This is the top cause when nvidia-smi works in your shell but Ollama reports no GPU — the process Ollama runs as is not the process you tested with.

The one-command test for Docker: run nvidia-smi inside the container.

docker exec -it <container-name> nvidia-smi

If that fails inside but works on the host, Ollama is not the problem. You need the NVIDIA Container Toolkit on the host and the GPU flag on the run command:

docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

Ollama's docs also call out a cgroup driver mismatch as a distinct Docker cause — the fix is adding "exec-opts": ["native.cgroupdriver=cgroupfs"] to /etc/docker/daemon.json and restarting Docker. If you run Ollama behind a web UI in Compose, the same flags have to be on the Ollama service specifically; see our Ollama + Open WebUI Docker setup.

The one-command test for systemd: check what the service user can see, not what you can see.

sudo -u ollama nvidia-smi

On AMD this is the usual failure. Ollama's docs are explicit that GPU access requires video and/or render group membership for /dev/kfd:

sudo usermod -aG video,render ollama
sudo systemctl restart ollama

Two smaller variants of the same class of problem, both documented: a noexec /tmp on hardened Linux boxes stops Ollama extracting its runtime libraries — set OLLAMA_TMPDIR to a writable, executable location. And if you run more than one card, a stale device filter can hide one of them; the multi-GPU specifics are in our Ollama multi-GPU setup guide.


Cause 6: Ollama Picked Vulkan Instead of CUDA

On a machine with both backends installed, check the startup log for which one was selected before changing anything. Ollama's hardware docs describe Vulkan as additional support on Windows and Linux, with device selection via GGML_VK_VISIBLE_DEVICES and an off switch at OLLAMA_VULKAN=0.

This matters because Vulkan and CUDA are not equivalent paths. A run can look "GPU accelerated" in ollama ps while behaving nothing like the CUDA numbers you read in a benchmark article — the backend that got selected is a variable most benchmark posts never state.

The one-command test:

journalctl -u ollama --no-pager | grep -i -E "vulkan|cuda|rocm" | head -20

If the log shows Vulkan on a machine with a working CUDA stack, force CUDA and compare:

OLLAMA_VULKAN=0 ollama serve

Then re-run the same prompt and compare tokens per second. If the numbers are identical, backend selection was not your problem — put the flag back and move on. We are deliberately not publishing a "Vulkan is X% slower" figure here. We do not own this hardware, and any single number would depend on the card, the driver and the model — your own before-and-after on the same prompt is the only measurement that means anything.


Cause 7: Something Is Explicitly Forcing CPU

Check for leftover environment variables before you rebuild anything — a stale CUDA_VISIBLE_DEVICES=-1 in a shell profile or a systemd drop-in produces a perfect 100% CPU with no error at all. Ollama's docs document -1 as the value that forces CPU-only mode for CUDA, with ROCR_VISIBLE_DEVICES doing the same job on AMD.

The one-command test:

# what the running server actually inherited (Linux)
sudo tr '\0' '\n' < /proc/$(pgrep -f 'ollama serve' | head -1)/environ | grep -i -E "cuda|rocr|ollama|vulkan|hsa"

That reads the environment of the live process, which is the only version that counts. Your interactive shell can be spotless while the systemd unit carries a two-year-old override. The usual suspects:

  • CUDA_VISIBLE_DEVICES=-1 or ROCR_VISIBLE_DEVICES=-1 — forces CPU
  • OLLAMA_LLM_LIBRARY=cpu_avx2 — pins the CPU backend and bypasses autodetection entirely. Ollama's troubleshooting docs list cpu_avx2 (best performance), cpu_avx and cpu (most compatible). It is a genuinely useful debugging tool and a genuinely common thing to leave behind after debugging
  • HSA_OVERRIDE_GFX_VERSION set to a wrong target on AMD — documented as the override for mapping unsupported GPUs onto a similar LLVM target, and it will happily map you onto something that does not work. Setting it on a card that is already supported is a common self-inflicted wound; see cause 8 for which cards need one and which do not
# find the systemd override that is doing it
systemctl cat ollama | grep -i environment

The WSL2 variant. If Ollama is installed inside WSL2, it needs the WSL CUDA driver path to work at all — test it with wsl nvidia-smi from PowerShell. And WSL2 has a second, sneakier failure: it can allow allocations to spill into shared system memory rather than failing, which reads as "GPU is being used" while running at CPU-like speed. If your WSL2 numbers are far below what the same card does natively, install Ollama on Windows directly and compare; our Ollama Windows installation guide covers the native path.


Cause 8: On AMD, Your gfx Target Is Not One Ollama Built For

If the log says amdgpu is not supported and names a gpu_type=, your driver is fine and your card is fine — Ollama simply did not ship a compiled code object for your chip's architecture. HSA_OVERRIDE_GFX_VERSION makes the runtime report a different architecture, which works when your chip is a configuration variant of a supported one and fails badly when it is not.

This is by far the biggest gap in most "Ollama not using GPU" guides, and the tracker shows why: a GitHub search of ollama/ollama on 2026-08-22 returns 966 issues and pull requests mentioning HSA_OVERRIDE_GFX_VERSION and 1,227 mentioning ROCm. The most-discussed hardware threads in the repository are AMD ones — #738 "AMD GPU & ROCm support" at 323 comments, #2453 "Add support for older AMD GPU gfx803, gfx802, gfx805 (e.g. Radeon RX 580, FirePro W7100)" at 222, PR #6282 "AMD integrated graphic on linux kernel 6.9.9+, GTT memory, loading freeze fix" at 195, and #2637 "Integrated AMD GPU support" at 171.

Step 1 — get your real gfx target. Do not guess it from the marketing name.

rocminfo | grep -i gfx
# or, if rocminfo is not installed
journalctl -u ollama --no-pager | grep -o 'gpu_type=gfx[0-9a-z]*'

Step 2 — look it up. Set an override only if the row says to.

Productgfx targetSet HSA_OVERRIDE_GFX_VERSION toWhy
RX 7900 XTX / XT / GREgfx1100nothingSupported by name; an override here only breaks a working card
RX 7800 XT / 7700 XTgfx1101nothingSupported by name
RX 7600 / 7600 XTgfx1102nothinggfx1102 code object ships
RX 6900 XT / 6800 XT / 6800gfx1030nothinggfx1030 is in Ollama's built list
RX 6700 XT / 6750 XT / 6700gfx103110.3.0RDNA 2 configuration variant of gfx1030 — same ISA
RX 6600 / 6600 XT / 6650 XTgfx103210.3.0Same reason
RX 6500 XT / 6400gfx103410.3.0Works, but 4GB caps you at small models anyway
Radeon 680M / 660M iGPUgfx103510.3.0RDNA 2 iGPU; see the GTT note below
Radeon 780M / 760M iGPUgfx110311.0.0 on older stacks; nothing on ROCm 7.xgfx1103 ships in current ROCm; the override is a legacy workaround
RX 5700 XT / 5700 / 5600 XTgfx1010no reliable valueRDNA 1 is a different ISA generation — use Vulkan instead
Radeon VII, Instinct MI50 / MI60gfx906override will not helpDropped from current ROCm; no gfx906 binary ships
RX 580 / 570 / 480, FirePro W7100gfx803override will not helpOut of ROCm support since ROCm 4.x; this is what #2453 is about

Product-to-target mappings are from AMD's ROCm compatibility matrix and the LLVM AMDGPU processor table; the built-target list is whatever your own log prints in the supported types:[...] bracket, which is the only list that is authoritative for the build you are running. Read that bracket before you set anything — it changes between Ollama releases.

Step 3 — set it where the server will actually see it. Exporting it in your shell does nothing for a systemd service:

sudo systemctl edit ollama
# add:
#   [Service]
#   Environment="HSA_OVERRIDE_GFX_VERSION=10.3.0"
sudo systemctl daemon-reload && sudo systemctl restart ollama

The value must be exactly three numbers — major.minor.stepping. HSA_OVERRIDE_GFX_VERSION=11.0 is not shorthand for 11.0.0; it is a parse error. On a machine with two different AMD targets, use the per-node form HSA_OVERRIDE_GFX_VERSION_0 / _1 so you do not convert a working card into a broken one. The full per-card table, the parsing rules and the failure modes when an override is wrong are in our ROCm unsupported-GPU and HSA_OVERRIDE_GFX_VERSION guide; the from-scratch install path is in the AMD ROCm local LLM setup guide.

Three AMD failure modes an override will not touch:

  1. Group membership. /dev/kfd needs video and/or render — that is cause 5, and on AMD it is the most common one of all.
  2. iGPU memory accounting. On APUs, Linux 6.9.9+ changed how GTT memory is reported, and Ollama has historically skipped integrated GPUs whose apparent VRAM is too small. PR #6282 documents the symptom — server hangs and out-of-memory on APUs — and lists the affected targets (gfx1103, gfx1037, gfx1035, gfx1033, gfx1036, gfx1151, gfx1152, gfx940, gfx90c). Note that PR was closed without being merged, so treat it as a description of the problem rather than a shipped fix; check your own version's behaviour. PR #5426 is a community write-up of the 780M path on Linux.
  3. A missing ROCm userspace. amdgpu detected, but no compatible rocm library found is not an override problem. Install the ROCm version Ollama's docs ask for, then re-read the log.

If none of that lands, Vulkan is the pragmatic fallback on AMD — it does not care about gfx code objects. That is cause 6 in reverse: instead of forcing CUDA off, you leave Vulkan on.


What This Page Does Not Cover

Being straight about the edges of this page. Everything here is sourced from Ollama's documentation, its issue tracker and AMD's published support matrices. We do not own this hardware and there are no first-hand benchmarks on this page — where a number appears, either the source is named or the arithmetic is shown so you can check it.

  • No performance figures. Not for Vulkan versus CUDA, not for one card against another. We can tell you how to see which backend was selected and how to force the other one; the tokens-per-second comparison has to be yours, on your prompt.
  • The compute-capability table is transcribed, not tested. It comes from Ollama's published hardware docs. If you are on a 5.x or 6.x card, treat the driver-570 requirement as the documented requirement and verify against the docs shipped with your Ollama version.
  • The AMD override table is a mapping, not a guarantee. Product-to-gfx mappings come from AMD and LLVM; whether a given override works on your build depends on the supported types:[...] list your own log prints. Where LLVM lists a target's retail names as TBA — gfx1032 and gfx1034 are both in that state — the mapping is the long-established one from rocminfo output, and we have flagged it rather than dressing it up as official.
  • Issue-tracker counts are point-in-time. The 152 / 592 / 966 / 1,227 figures are GitHub search results for ollama/ollama on 2026-08-22 and will drift.

If a step here does not match what you see, the version you are running is the first thing to check. ollama --version — and if it is not v0.32.x, expect different log lines.


Sources


FAQ

🎯
AI Learning Path

Ollama’s running. Here’s what to build with it.

Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.

Or own it for life — Lifetime $149 $599, pay once
Once your hardware is sorted

Stop piecing Ollama together from blog posts

Ollama Mastery is 15 chapters end to end — install, model choice, Modelfiles, GPU offload, the API, and the 20 errors that actually happen. Plus 24 more courses.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsOllamaGPUCUDAROCmAMDTroubleshootingVRAMKV CacheVulkan

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on Ollama
See the full Best Ollama Models 2026 guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

How do I check whether Ollama is actually using my GPU?

Load a model, then run `ollama ps` in a second terminal. The PROCESSOR column is the answer: "100% GPU" means the whole model is on the card, "100% CPU" means none of it is, and a split like "48%/52% CPU/GPU" means Ollama loaded the model partially onto the GPU and partially into system memory. Ollama's FAQ documents exactly these three states. nvidia-smi is a weaker test — it shows the process and the VRAM allocation but not the layer split, and utilisation samples can read near zero between token batches even on a healthy run.

Why does Ollama use the CPU when my model is smaller than my VRAM?

Because the model weights are not the only thing that has to fit. The KV cache is allocated separately and scales linearly with context length. At Ollama's default 4096-token context an 8B-class model with 36 layers and 8 KV heads needs roughly 0.56 GB of f16 KV cache; raise OLLAMA_CONTEXT_LENGTH to 32768 and the same model needs about 4.5 GB. Add the compute buffers and a desktop compositor already holding 1-2 GB, and a "4 GB model on a 12 GB card" can genuinely run out of room. `ollama ps` will show a split rather than 100% GPU when this happens.

What NVIDIA driver version does Ollama require?

Ollama's GPU documentation states driver 550 or newer as the general floor, and driver 570 or newer specifically for GPUs with compute capability 5.0 through 6.2 — which is the Maxwell and Pascal generation, including the GTX 900 and GTX 10-series. Check with `nvidia-smi --query-gpu=driver_version --format=csv,noheader`. A driver below the floor is one of the few causes that produces a completely silent fall back to CPU.

Does Ollama still support Pascal cards like the GTX 1080 Ti?

Yes, per the current hardware documentation. Ollama lists compute capability 5.0 as the minimum, and the supported table includes 6.1 (GTX 10-series, TITAN Xp), 6.0 (Tesla P100), 5.2 (GTX 900-series) and 5.0 (GTX 750-series). The catch is the driver: those 5.0-6.2 cards need driver 570+, so a Pascal box on an older long-term-support driver will run on CPU while nvidia-smi looks perfectly healthy. If you are on an older card, verify against Ollama's own docs for the version you are running rather than trusting a forum post — the support matrix has moved before.

Ollama picked Vulkan instead of CUDA. How do I force CUDA?

Set OLLAMA_VULKAN=0 and restart the server. Ollama's hardware docs describe Vulkan as additional support on Windows and Linux, controlled with GGML_VK_VISIBLE_DEVICES and disabled with OLLAMA_VULKAN=0. On a machine where both backends are installed, check the server log at startup to see which one was selected before you change anything — if the log shows CUDA, Vulkan is not your problem and disabling it will not help.

What does "no compatible GPUs were discovered" mean in the Ollama log?

It means GPU discovery ran at server start and finished with an empty list — Ollama never even attempted to use a card, so nothing later in the run will fix it. The line is logged as `level=INFO source=gpu.go:<line> msg="no compatible GPUs were discovered"`, and the INFO level is the trap: people grep their logs for "error" or "fail" and never see it. On AMD it is usually preceded by a WARN from `amd_linux.go` naming your gfx target, which tells you whether the cause is the driver, the container, or an unsupported architecture.

My AMD card shows "amdgpu is not supported". Does HSA_OVERRIDE_GFX_VERSION fix it?

Sometimes. That WARN line prints your real gfx target and the list of targets Ollama shipped code objects for, e.g. `gpu_type=gfx1031` against `supported types:[gfx1030 gfx1100 gfx1101 gfx1102 gfx900 gfx906 gfx908 gfx90a gfx940 gfx941 gfx942]`. HSA_OVERRIDE_GFX_VERSION only changes the architecture string the runtime reports — it cannot add a code object that was never built. It works when your chip is a configuration variant of a supported target in the same ISA family, which is why RDNA 2 parts like gfx1031, gfx1032 and gfx1034 map onto 10.3.0. It does not work across ISA generations, and it does not resurrect gfx803 (RX 580) or gfx906 (Radeon VII, MI50), where no matching binary ships at all.

Why does Ollama see no GPU inside Docker when nvidia-smi works on the host?

The container was almost certainly started without GPU access. The container needs the NVIDIA Container Toolkit installed on the host and a `--gpus=all` flag on `docker run`. The fastest proof is to run nvidia-smi inside the container itself: if it fails there but works on the host, the problem is the container runtime, not Ollama. Ollama's troubleshooting docs also flag a cgroup driver mismatch as a cause, fixed by adding "exec-opts": ["native.cgroupdriver=cgroupfs"] to /etc/docker/daemon.json.

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks with GPU passthrough already wired up — skip the --gpus debugging entirely. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: September 13, 2026🔄 Last Updated: September 13, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Ollama’s running. Here’s what to build with it.

Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators