Mac Out of Memory on Local LLMs: The Metal VRAM Fix
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Got the hardware sorted? Now build on it. You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.
Short answer: macOS will not let Metal wire all of your unified memory, the cap is exposed as recommendedMaxWorkingSetSize, and you raise it with sudo sysctl iogpu.wired_limit_mb=<MB> — but in most cases raising it is the wrong fix. If a model needs more than your machine has, the cap is not what is hurting you; swap is. Read your real limit first (one command, below), then decide between raising the cap by a few gigabytes and dropping one quantisation level.
There is also a second, newer problem that looks identical from the outside: on Ollama's MLX engine, ollama stop can report success, clear ollama ps, and leave the runner process holding gigabytes anyway. That one is not a settings problem and no sysctl fixes it.
Read Your Real Limit First
Do not take a percentage figure from a blog post — including this one. Print the number your own machine reports. The value that matters is Metal's recommendedMaxWorkingSetSize, and three tools will show it to you.
If you have MLX installed:
python3 -c "import mlx.core as mx; d = mx.metal.device_info(); \
print(d['memory_size'] / 1e9, 'GB total'); \
print(d['max_recommended_working_set_size'] / 1e9, 'GB wireable')"
Apple's MLX documentation describes max_recommended_working_set_size as the system wired limit and memory_size as total memory, and uses exactly this pair to explain why set_wired_limit() will refuse a value above the system limit.
If you use llama.cpp or anything built on it, you already have the number in your logs — it prints on every Metal init:
ggml_metal_init: recommendedMaxWorkingSetSize = 51539.61 MB
llama.cpp reads it straight from MTLDevice.recommendedMaxWorkingSetSize and, when a model overruns it, logs warning: current allocated size is greater than the recommended max working set size. That warning line in your terminal is the single most useful signal on this whole page: it tells you the run is over budget before the beachball starts.
And to see whether an override is currently in place:
sysctl iogpu.wired_limit_mb
That reports the override value, not the effective driver default — so treat the Metal-reported number as authoritative and this one as "has somebody already changed this?".
Write both numbers down. Total RAM minus wireable memory is your working headroom for macOS itself, and every decision below is made against those two figures.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
Symptom → Cause → Fix
Ordered by how often each one is actually the culprit, not by how interesting it is. Match the symptom, then jump to the section.
| What you see | Real cause | Fix |
|---|---|---|
| Memory pressure red, UI stutters, tok/s collapses partway through a long reply | KV cache growth — the model fit, the conversation didn't | Shorter context, or a smaller quant. Not the wired limit |
| Runner killed on load, before a single token | Model + cache genuinely exceeds wireable memory | Drop one quantisation level; see quantization explained |
RAM still held after ollama stop, ollama ps empty | Ollama MLX runner subprocess staying resident (open bug) | Kill the runner |
| Loads and runs, but a couple of GB short of stable on a 64GB+ machine | Genuine wired-limit headroom problem | Raise the limit — this is the case it exists for |
MPS backend out of memory ... max allowed | PyTorch MPS allocator watermark, not macOS | Watermark ratios |
| Enormous single allocation request on an image input | MLX vision runner buffer-size bug | Downscale the image; see below |
The ranking matters. The reader who lands here from iogpu.wired_limit_mb usually arrives already convinced the sysctl is the answer, and for most of them it is the least likely of the six.
Why the KV cache is first on that list
A 27B model at Q4 is roughly 16-17GB of weights. That is the number people plan around, and it is the wrong number, because the KV cache grows with context length and lives in the same pool. A model that loads with 4GB to spare at an empty prompt can be in swap forty turns later without anything about your configuration changing. If your collapse happens during a session rather than at load, stop tuning memory limits and shorten the context window instead.
Raising the Wired Limit Safely
The command is one line, and the constraint that keeps your Mac alive is one sentence from Apple's own docs: the wired limit "should remain strictly less than the total memory size."
sudo sysctl iogpu.wired_limit_mb=49152 # 48GB, on a 64GB machine
This is the command Apple's MLX documentation names for increasing the system wired limit. It takes effect immediately and is gone after a reboot. People who want it persistent add the same key to /etc/sysctl.conf — that is the approach you will see in llama.cpp threads such as issue #13361, which uses iogpu.wired_limit_mb=122880 on a large-memory machine. Treat persistence as the riskier choice: a value that freezes your Mac is much easier to recover from when a reboot clears it.
Real values from public issue threads, for calibration on what practitioners actually set:
| Machine context | Value seen | Source |
|---|---|---|
| 24GB-class, MoE offload experiment | iogpu.wired_limit_mb=22000 (~21.5GB) | llama.cpp issue #19825 (Feb 2026) |
| 128GB, two-node Metal RPC setup | iogpu.wired_limit_mb=117760 (115GB) | llama.cpp issue #26384 (Jul 2026) |
| Large-memory, persisted via sysctl.conf | iogpu.wired_limit_mb=122880 (120GB) | llama.cpp issue #13361 |
Note the pattern in all three: roughly 90% of installed RAM, never 100%. That is the whole heuristic.
Headroom arithmetic — this is subtraction, not a benchmark, and we say so plainly:
| Installed RAM | Leave for macOS | Resulting ceiling | Honest read |
|---|---|---|---|
| 16GB | ~4GB | ~12GB | Do not touch the sysctl. 8B Q4 is your lane |
| 24GB | ~5GB | ~19GB | Marginal gain. 14B Q4 comfortable, 27B Q4 is a fight |
| 32GB | ~6GB | ~26GB | Worth doing to stabilise a 27B Q4 |
| 64GB | ~8GB | ~56GB | The sweet spot for this tweak |
| 128GB | ~10GB | ~118GB | Matches what large-memory users actually set |
The "leave for macOS" column is a judgement call, not a measurement. Browsers, Slack and an IDE will happily eat 8GB between them; if that is your desktop, be more generous than the table.
The failure mode to respect: over-raise this and you do not get a polite error, you get a machine that stops redrawing. macOS has nothing left to wire for the window server. On a laptop, with unsaved work, at 11pm. Move in increments of a few GB and reboot to clear if it goes wrong.
The Ollama MLX Runner Problem
If ollama ps is empty and Activity Monitor still shows gigabytes gone, you are not misreading it — this is an open Ollama bug on the MLX engine, filed 16 August 2026.
Issue ollama/ollama #17792 is titled, verbatim: "ollama stop reports success and clears ollama ps, but the MLX runner subprocess stays resident and keeps holding RAM until manually killed." It was reported against Ollama 0.32.9 and 0.32.13, across several models, with retained memory ranging from roughly 1GB to about 20GB. It is open as of 18 August 2026.
It is not the only MLX-engine memory report in that window, which is why we treat it as a class of problem rather than a one-off:
| Issue | Opened | What it says |
|---|---|---|
| #17792 | 2026-08-16 | Runner stays resident after ollama stop; ollama ps empty |
| #17804 | 2026-08-16 | MLX vision runner requests ~125GB Metal buffer on high-resolution images |
| #17783 | 2026-08-15 | Model size in memory grows across calls (gemma4:31b-mlx) |
| #17829 | 2026-08-17 | No prefix caching between requests; reported 26 tok/s falling to 3 tok/s as history grows |
| #16698 | 2026-06-13 | KV cache not released between requests (v0.30.8) — closed |
All five are from Ollama's own tracker; #16698 is closed, the rest were open when we checked.
What to actually do
Confirm the process is really still there before blaming anything else:
ollama ps # will likely show nothing
ps aux | grep -i "ollama.*runner" | grep -v grep
If a runner is listed, kill that PID directly. Restarting the Ollama service clears it too, and is the safer move if you are not certain which PID is which:
kill <pid> # the specific runner
# or, blunt but reliable:
pkill -f "ollama.*runner"
Then re-check with Activity Monitor rather than ollama ps, since ollama ps is precisely the reading that was wrong.
Prevention, not cure. Ollama's FAQ documents that models stay in memory five minutes by default, that OLLAMA_KEEP_ALIVE sets that globally, and that a keep_alive of 0 on an API request unloads immediately. If you are cycling between large models on a Mac, a short keep-alive is worth setting regardless of this bug — you stop paying for a model you finished with ten minutes ago.
Version numbers matter here. The MLX engine's behaviour has been moving within the 0.32.x series, so check your own version with ollama --version and check whether #17792 has closed before assuming the symptom is still current.
Which runner is even serving you?
Tags carry it explicitly — a model pulled as gemma4:31b-mlx is on the MLX path, and a plain GGUF tag is not. If the tag does not tell you, the loading logs will: the MLX runner and the GGML/Metal runner announce themselves differently, and only the llama.cpp path prints the recommendedMaxWorkingSetSize line quoted earlier. If you see that line, you are on GGUF.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
MPS Errors: PyTorch and ComfyUI
The exact string is MPS backend out of memory (MPS allocated: ..., other allocations: ..., max allowed: ...). Tried to allocate ... on ... pool. Use PYTORCH_MPS_HIGH_WATERMARK_RATIO=0.0 to disable upper limit for memory allocations (may cause system failure). That is from PyTorch's MPS allocator source, and the parenthetical warning is PyTorch's, not ours.
The part almost everybody gets wrong: the watermark ratios are multiples of recommendedMaxWorkingSetSize, not fractions of your RAM. PyTorch's MPSAllocator.h sets them out directly:
| Constant | Default | Meaning |
|---|---|---|
default_high_watermark_ratio | 1.7 | Allocation ceiling, as a multiple of the recommended working set |
default_high_watermark_upper_bound | 2.0 | The highest ratio you are allowed to set |
default_low_watermark_ratio | 1.4 | Above this, the allocator starts collecting garbage |
So the allocator is already permitted to run 70% past what Metal recommends before it raises anything, and the source comment explains why — on unified memory you can allocate beyond the recommended working set. Advice telling you to "set it to 0.8" is advice from someone who read the name and not the header.
Consequences worth knowing:
- Setting
PYTORCH_MPS_HIGH_WATERMARK_RATIO=0.0does not raise a ceiling, it removes it. Your process will then allocate until macOS makes the decision for you, which is exactly the "system failure" the error text warns about. - Because the default is already 1.7, an MPS OOM means you were at roughly 1.7× the recommended working set. You are not a tweak away from fitting. Reduce resolution, batch size or model size.
torch.mps.empty_cache()between stages releases cached-but-unused blocks and is the sane first move in a notebook or a long ComfyUI session.
For image workflows specifically, the same over-budget diagnosis applies as on any other platform — our ComfyUI out-of-memory guide covers the node-graph side of it.
What We Could Not Test
We are not going to invent a benchmark table. This page is built from vendor documentation and from the vendors' own public issue trackers, both checked on 18 August 2026, and it is honest about the seams:
- The default wired-limit percentage. You will see "75%" quoted widely. We could not confirm a current, first-party figure for it, and it is machine-dependent anyway — which is why the first section tells you to print your own number rather than trust a percentage.
- Side-by-side tok/s at default versus raised limits. That requires sustained measurement on several Apple Silicon RAM tiers, and we do not publish numbers we have not measured. The public issue values in the table above are labelled as what other people set, not as our results.
- Whether #17792 is fixed by the time you read this. Check the issue and your
ollama --versionbefore assuming.
If you are still deciding what to buy rather than what to fix, the Apple Silicon AI calculator sizes models against unified memory, and the Apple Silicon buying guide covers the tiers. If you already have the machine, Mac local AI setup and running Llama on a Mac are the practical starting points, and the Ollama model RAM/VRAM table will tell you which tags fit before you download 17GB to find out.
The Short Version
- Print
recommendedMaxWorkingSetSizebefore changing anything. One MLX command, or one line already in your llama.cpp logs. - Collapse mid-conversation is KV cache, not the wired limit. Shorten context first.
- Raise
iogpu.wired_limit_mbonly on 32GB+ machines, in small steps, staying well under total RAM. Apple's own docs say strictly less than total memory; practitioners sit around 90%. - If
ollama stopdid not free the RAM, kill the runner. Open bug #17792, seen on 0.32.9 and 0.32.13. - MPS watermarks are multiples of the recommended working set (default 1.7), not fractions of RAM. An MPS OOM means shrink the workload, not tune the ratio.
The uncomfortable conclusion for most readers: the model you want does not fit, and one quantisation level down is a better trade than any setting on this page.
Sources
- MLX documentation —
mlx.core.set_wired_limit— theiogpu.wired_limit_mbcommand, the strictly-less-than-total constraint, anddevice_info() - llama.cpp Metal device backend —
recommendedMaxWorkingSetSizelogging and the over-allocation warning; issues #13361, #19825, #26384 for real-world sysctl values - Ollama issue tracker — #17792, #17804, #17783, #17829, #16698 (MLX-engine memory behaviour, Aug 2026)
- Ollama FAQ — five-minute default keep-alive,
OLLAMA_KEEP_ALIVE,ollama stop - PyTorch
MPSAllocator— the exact OOM string and the 1.7 / 1.4 / 2.0 watermark defaults
All checked 18 August 2026.
FAQ
Got the hardware sorted? Now build on it.
You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.
Decide before you spend a thousand pounds
The AI Hardware course sizes your build properly — VRAM ladder, real bottlenecks, budget builds — and Pick the Right Model tells you what to run on it.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARLocal AI Hardware Requirements (2026): Complete Guide
- AI Hardware Requirements: CPU, GPU and RAM for Beginners
- AI RAM Requirements 2026: How Much for 7B, 13B, 70B Models?
- AI Server Build Under $1,500: Parts List and What Fits
- AMD GPU Not Supported by ROCm? HSA_OVERRIDE Values
- AMD MI50 32GB for Local LLMs: The Used VRAM King, Honestly
- AMD Ryzen AI Max+ 395 (Strix Halo) for Local AI 2026
- Apple M4 for Local AI: Mac Studio + MacBook Guide (2026)
- Benchmark Your Local AI Setup: tok/s, TTFT, VRAM
- Best GPU for AI Video Generation: By VRAM Tier (2026)
Comments (0)
No comments yet. Be the first to share your thoughts!