★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Mac

Mac Out of Memory on Local LLMs: The Metal VRAM Fix

September 27, 2026
11 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Got the hardware sorted? Now build on it. You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Start free
Or own it for life — Lifetime $149, pay once

Short answer: macOS will not let Metal wire all of your unified memory, the cap is exposed as recommendedMaxWorkingSetSize, and you raise it with sudo sysctl iogpu.wired_limit_mb=<MB> — but in most cases raising it is the wrong fix. If a model needs more than your machine has, the cap is not what is hurting you; swap is. Read your real limit first (one command, below), then decide between raising the cap by a few gigabytes and dropping one quantisation level.

There is also a second, newer problem that looks identical from the outside: on Ollama's MLX engine, ollama stop can report success, clear ollama ps, and leave the runner process holding gigabytes anyway. That one is not a settings problem and no sysctl fixes it.


Read Your Real Limit First

Do not take a percentage figure from a blog post — including this one. Print the number your own machine reports. The value that matters is Metal's recommendedMaxWorkingSetSize, and three tools will show it to you.

If you have MLX installed:

python3 -c "import mlx.core as mx; d = mx.metal.device_info(); \
print(d['memory_size'] / 1e9, 'GB total'); \
print(d['max_recommended_working_set_size'] / 1e9, 'GB wireable')"

Apple's MLX documentation describes max_recommended_working_set_size as the system wired limit and memory_size as total memory, and uses exactly this pair to explain why set_wired_limit() will refuse a value above the system limit.

If you use llama.cpp or anything built on it, you already have the number in your logs — it prints on every Metal init:

ggml_metal_init: recommendedMaxWorkingSetSize  = 51539.61 MB

llama.cpp reads it straight from MTLDevice.recommendedMaxWorkingSetSize and, when a model overruns it, logs warning: current allocated size is greater than the recommended max working set size. That warning line in your terminal is the single most useful signal on this whole page: it tells you the run is over budget before the beachball starts.

And to see whether an override is currently in place:

sysctl iogpu.wired_limit_mb

That reports the override value, not the effective driver default — so treat the Metal-reported number as authoritative and this one as "has somebody already changed this?".

Write both numbers down. Total RAM minus wireable memory is your working headroom for macOS itself, and every decision below is made against those two figures.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

Symptom → Cause → Fix

Ordered by how often each one is actually the culprit, not by how interesting it is. Match the symptom, then jump to the section.

What you seeReal causeFix
Memory pressure red, UI stutters, tok/s collapses partway through a long replyKV cache growth — the model fit, the conversation didn'tShorter context, or a smaller quant. Not the wired limit
Runner killed on load, before a single tokenModel + cache genuinely exceeds wireable memoryDrop one quantisation level; see quantization explained
RAM still held after ollama stop, ollama ps emptyOllama MLX runner subprocess staying resident (open bug)Kill the runner
Loads and runs, but a couple of GB short of stable on a 64GB+ machineGenuine wired-limit headroom problemRaise the limit — this is the case it exists for
MPS backend out of memory ... max allowedPyTorch MPS allocator watermark, not macOSWatermark ratios
Enormous single allocation request on an image inputMLX vision runner buffer-size bugDownscale the image; see below

The ranking matters. The reader who lands here from iogpu.wired_limit_mb usually arrives already convinced the sysctl is the answer, and for most of them it is the least likely of the six.

Why the KV cache is first on that list

A 27B model at Q4 is roughly 16-17GB of weights. That is the number people plan around, and it is the wrong number, because the KV cache grows with context length and lives in the same pool. A model that loads with 4GB to spare at an empty prompt can be in swap forty turns later without anything about your configuration changing. If your collapse happens during a session rather than at load, stop tuning memory limits and shorten the context window instead.


Raising the Wired Limit Safely

The command is one line, and the constraint that keeps your Mac alive is one sentence from Apple's own docs: the wired limit "should remain strictly less than the total memory size."

sudo sysctl iogpu.wired_limit_mb=49152     # 48GB, on a 64GB machine

This is the command Apple's MLX documentation names for increasing the system wired limit. It takes effect immediately and is gone after a reboot. People who want it persistent add the same key to /etc/sysctl.conf — that is the approach you will see in llama.cpp threads such as issue #13361, which uses iogpu.wired_limit_mb=122880 on a large-memory machine. Treat persistence as the riskier choice: a value that freezes your Mac is much easier to recover from when a reboot clears it.

Real values from public issue threads, for calibration on what practitioners actually set:

Machine contextValue seenSource
24GB-class, MoE offload experimentiogpu.wired_limit_mb=22000 (~21.5GB)llama.cpp issue #19825 (Feb 2026)
128GB, two-node Metal RPC setupiogpu.wired_limit_mb=117760 (115GB)llama.cpp issue #26384 (Jul 2026)
Large-memory, persisted via sysctl.confiogpu.wired_limit_mb=122880 (120GB)llama.cpp issue #13361

Note the pattern in all three: roughly 90% of installed RAM, never 100%. That is the whole heuristic.

Headroom arithmetic — this is subtraction, not a benchmark, and we say so plainly:

Installed RAMLeave for macOSResulting ceilingHonest read
16GB~4GB~12GBDo not touch the sysctl. 8B Q4 is your lane
24GB~5GB~19GBMarginal gain. 14B Q4 comfortable, 27B Q4 is a fight
32GB~6GB~26GBWorth doing to stabilise a 27B Q4
64GB~8GB~56GBThe sweet spot for this tweak
128GB~10GB~118GBMatches what large-memory users actually set

The "leave for macOS" column is a judgement call, not a measurement. Browsers, Slack and an IDE will happily eat 8GB between them; if that is your desktop, be more generous than the table.

The failure mode to respect: over-raise this and you do not get a polite error, you get a machine that stops redrawing. macOS has nothing left to wire for the window server. On a laptop, with unsaved work, at 11pm. Move in increments of a few GB and reboot to clear if it goes wrong.


The Ollama MLX Runner Problem

If ollama ps is empty and Activity Monitor still shows gigabytes gone, you are not misreading it — this is an open Ollama bug on the MLX engine, filed 16 August 2026.

Issue ollama/ollama #17792 is titled, verbatim: "ollama stop reports success and clears ollama ps, but the MLX runner subprocess stays resident and keeps holding RAM until manually killed." It was reported against Ollama 0.32.9 and 0.32.13, across several models, with retained memory ranging from roughly 1GB to about 20GB. It is open as of 18 August 2026.

It is not the only MLX-engine memory report in that window, which is why we treat it as a class of problem rather than a one-off:

IssueOpenedWhat it says
#177922026-08-16Runner stays resident after ollama stop; ollama ps empty
#178042026-08-16MLX vision runner requests ~125GB Metal buffer on high-resolution images
#177832026-08-15Model size in memory grows across calls (gemma4:31b-mlx)
#178292026-08-17No prefix caching between requests; reported 26 tok/s falling to 3 tok/s as history grows
#166982026-06-13KV cache not released between requests (v0.30.8) — closed

All five are from Ollama's own tracker; #16698 is closed, the rest were open when we checked.

What to actually do

Confirm the process is really still there before blaming anything else:

ollama ps                       # will likely show nothing
ps aux | grep -i "ollama.*runner" | grep -v grep

If a runner is listed, kill that PID directly. Restarting the Ollama service clears it too, and is the safer move if you are not certain which PID is which:

kill <pid>                      # the specific runner
# or, blunt but reliable:
pkill -f "ollama.*runner"

Then re-check with Activity Monitor rather than ollama ps, since ollama ps is precisely the reading that was wrong.

Prevention, not cure. Ollama's FAQ documents that models stay in memory five minutes by default, that OLLAMA_KEEP_ALIVE sets that globally, and that a keep_alive of 0 on an API request unloads immediately. If you are cycling between large models on a Mac, a short keep-alive is worth setting regardless of this bug — you stop paying for a model you finished with ten minutes ago.

Version numbers matter here. The MLX engine's behaviour has been moving within the 0.32.x series, so check your own version with ollama --version and check whether #17792 has closed before assuming the symptom is still current.

Which runner is even serving you?

Tags carry it explicitly — a model pulled as gemma4:31b-mlx is on the MLX path, and a plain GGUF tag is not. If the tag does not tell you, the loading logs will: the MLX runner and the GGML/Metal runner announce themselves differently, and only the llama.cpp path prints the recommendedMaxWorkingSetSize line quoted earlier. If you see that line, you are on GGUF.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

MPS Errors: PyTorch and ComfyUI

The exact string is MPS backend out of memory (MPS allocated: ..., other allocations: ..., max allowed: ...). Tried to allocate ... on ... pool. Use PYTORCH_MPS_HIGH_WATERMARK_RATIO=0.0 to disable upper limit for memory allocations (may cause system failure). That is from PyTorch's MPS allocator source, and the parenthetical warning is PyTorch's, not ours.

The part almost everybody gets wrong: the watermark ratios are multiples of recommendedMaxWorkingSetSize, not fractions of your RAM. PyTorch's MPSAllocator.h sets them out directly:

ConstantDefaultMeaning
default_high_watermark_ratio1.7Allocation ceiling, as a multiple of the recommended working set
default_high_watermark_upper_bound2.0The highest ratio you are allowed to set
default_low_watermark_ratio1.4Above this, the allocator starts collecting garbage

So the allocator is already permitted to run 70% past what Metal recommends before it raises anything, and the source comment explains why — on unified memory you can allocate beyond the recommended working set. Advice telling you to "set it to 0.8" is advice from someone who read the name and not the header.

Consequences worth knowing:

  • Setting PYTORCH_MPS_HIGH_WATERMARK_RATIO=0.0 does not raise a ceiling, it removes it. Your process will then allocate until macOS makes the decision for you, which is exactly the "system failure" the error text warns about.
  • Because the default is already 1.7, an MPS OOM means you were at roughly 1.7× the recommended working set. You are not a tweak away from fitting. Reduce resolution, batch size or model size.
  • torch.mps.empty_cache() between stages releases cached-but-unused blocks and is the sane first move in a notebook or a long ComfyUI session.

For image workflows specifically, the same over-budget diagnosis applies as on any other platform — our ComfyUI out-of-memory guide covers the node-graph side of it.


What We Could Not Test

We are not going to invent a benchmark table. This page is built from vendor documentation and from the vendors' own public issue trackers, both checked on 18 August 2026, and it is honest about the seams:

  • The default wired-limit percentage. You will see "75%" quoted widely. We could not confirm a current, first-party figure for it, and it is machine-dependent anyway — which is why the first section tells you to print your own number rather than trust a percentage.
  • Side-by-side tok/s at default versus raised limits. That requires sustained measurement on several Apple Silicon RAM tiers, and we do not publish numbers we have not measured. The public issue values in the table above are labelled as what other people set, not as our results.
  • Whether #17792 is fixed by the time you read this. Check the issue and your ollama --version before assuming.

If you are still deciding what to buy rather than what to fix, the Apple Silicon AI calculator sizes models against unified memory, and the Apple Silicon buying guide covers the tiers. If you already have the machine, Mac local AI setup and running Llama on a Mac are the practical starting points, and the Ollama model RAM/VRAM table will tell you which tags fit before you download 17GB to find out.


The Short Version

  1. Print recommendedMaxWorkingSetSize before changing anything. One MLX command, or one line already in your llama.cpp logs.
  2. Collapse mid-conversation is KV cache, not the wired limit. Shorten context first.
  3. Raise iogpu.wired_limit_mb only on 32GB+ machines, in small steps, staying well under total RAM. Apple's own docs say strictly less than total memory; practitioners sit around 90%.
  4. If ollama stop did not free the RAM, kill the runner. Open bug #17792, seen on 0.32.9 and 0.32.13.
  5. MPS watermarks are multiples of the recommended working set (default 1.7), not fractions of RAM. An MPS OOM means shrink the workload, not tune the ratio.

The uncomfortable conclusion for most readers: the model you want does not fit, and one quantisation level down is a better trade than any setting on this page.


Sources

All checked 18 August 2026.


FAQ

🎯
AI Learning Path

Got the hardware sorted? Now build on it.

You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Once your hardware is sorted

Decide before you spend a thousand pounds

The AI Hardware course sizes your build properly — VRAM ladder, real bottlenecks, budget builds — and Pick the Right Model tells you what to run on it.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsApple SiliconMetalMLXOllamaMemory PressureTroubleshooting

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on Local AI Hardware
See the full AI Hardware Guide 2026 guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

How much of my Mac's unified memory can the GPU actually use?

Less than all of it, and the exact number is machine-specific — so read it rather than trusting a blog figure. Metal exposes it as recommendedMaxWorkingSetSize. MLX surfaces the same value through mx.metal.device_info() as "max_recommended_working_set_size" alongside "memory_size" (your total RAM), and llama.cpp prints "recommendedMaxWorkingSetSize = ... MB" in its Metal init log every time you load a model. Compare those two numbers on your own machine and you have your real ceiling in about ten seconds.

What does sudo sysctl iogpu.wired_limit_mb actually do?

It raises the system-wide cap on how much memory can be wired for GPU use. Apple's MLX documentation points at exactly this command — "sudo sysctl iogpu.wired_limit_mb=<size_in_megabytes>" — as the way to increase the system wired limit, and warns that the wired limit "should remain strictly less than the total memory size." MLX will raise an error if you ask it for a wired limit above the system value. The sysctl is a runtime setting and does not survive a reboot on its own.

Is raising the wired limit the fix for memory pressure?

Usually not. Raising it removes a safety margin; it does not create memory. If a Q4 27B model needs ~17GB of weights plus KV cache on a 24GB machine, no sysctl makes that comfortable — you are asking macOS to wire almost everything and leave nothing for the window server, Safari and Spotlight, which is how people hard-freeze their Macs. The wired limit is the fix when you are a couple of gigabytes short on a large-RAM machine. A smaller quantisation is the fix the rest of the time.

Why is my Mac still holding gigabytes after ollama stop?

On the MLX engine this is a known, currently open bug. ollama/ollama issue #17792, opened 16 August 2026, is titled "ollama stop reports success and clears ollama ps, but the MLX runner subprocess stays resident and keeps holding RAM until manually killed" — reproduced on Ollama 0.32.9 and 0.32.13, with roughly 1GB to about 20GB retained depending on the model. Check for a leftover runner process and kill it; ollama ps will already show nothing.

What is PYTORCH_MPS_HIGH_WATERMARK_RATIO and what is the default?

It is the PyTorch MPS allocator's ceiling, expressed as a multiple of recommendedMaxWorkingSetSize — not a fraction of your RAM, which is where most advice on this goes wrong. In PyTorch's MPSAllocator header the defaults are default_high_watermark_ratio = 1.7, default_low_watermark_ratio = 1.4 and an upper bound of 2.0. So the allocator is already allowed to go 70% past the recommended working set before it throws. The error text itself suggests PYTORCH_MPS_HIGH_WATERMARK_RATIO=0.0 to remove the limit, and says in the same sentence that doing so "may cause system failure."

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks for local models — keep a second machine serving while your Mac gets its memory back. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: September 27, 2026🔄 Last Updated: September 27, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators