Image Generation on a Mac: M4, M4 Pro and M4 Max
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Generating images locally? Take it further. From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.
The honest number: Apple's own Core ML benchmark lists SDXL base at 1024x1024 and 20 steps at 37 seconds on a MacBook Pro M2 Max and 20 seconds on a Mac Studio M2 Ultra. Nobody publishes an equivalent measured table for M4, M4 Pro or M4 Max — including us — so anyone quoting you an exact M4 second-count is guessing. What is not a guess: on any Apple Silicon Mac, a large SDXL-class image lands in the tens of seconds, and your model choice moves that number far more than your chip tier does.
That last point is the one worth reading this page for. A 6B model that needs 8 sampling steps and a 12B model that needs 28 are not in the same universe of work, and no amount of M4 Max costs its way out of the difference. If you have a 16GB M4 Air and you have been told local image generation is off the table, it isn't — you just need to stop trying to run the model the CUDA tutorials assume.
Landed here with a traceback rather than a stopwatch? This page is about how long a Mac takes. If your run is failing outright — not currently implemented for the MPS device, float64, MPS backend out of memory — skip to When it errors instead of being slow, which triages the string and hands you to the fixes.
What Is Actually Measured on Apple Silicon
The only first-party seconds-per-image table for Mac diffusion is in Apple's ml-stable-diffusion repository, and it stops at the M2 generation. That is the whole state of the public record. Here is what Apple publishes, taken from their benchmark tables:
| Device | Model | Resolution / steps | Compute unit | End-to-end latency |
|---|---|---|---|---|
| Mac Studio M2 Ultra | SDXL base | 1024x1024, 20 steps | CPU_AND_GPU | 20 s |
| MacBook Pro M2 Max | SDXL base | 1024x1024, 20 steps | CPU_AND_GPU | 37 s |
| iPad Pro M2 | SDXL base | 768x768, 20 steps | CPU_AND_NE | 27 s |
| iPad Pro M2 | SD 2.1 base | 512x512, 20 steps | CPU_AND_NE | 7.0 s |
| iPhone 15 Pro Max | SDXL base | 768x768, 20 steps | CPU_AND_NE | 31 s |
| iPhone 14 Pro Max | SD 2.1 base | 512x512, 20 steps | CPU_AND_NE | 7.9 s |
Source: Apple, ml-stable-diffusion benchmark tables. Apple notes these are "the median latency value across 5 back-to-back end-to-end executions."
Two things fall straight out of that table. First, the spread between a laptop chip and the biggest desktop chip on the same job is under 2x (37 s to 20 s) — real, but not the difference between "usable" and "not". Second, resolution and step count swing the result harder than the silicon: the same M2 hardware family goes from 37 seconds to 7 seconds when the job drops from SDXL at 1024x1024 to SD 2.1 at 512x512.
We have not run these workloads on M4 hardware. Rather than extrapolate a number and dress it up as a measurement, the Benchmark Your Own Mac section below gives you a repeatable recipe that takes about ten minutes.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
Steps Beat Chips: The Arithmetic That Decides Your Wait
A diffusion image costs roughly (model weights moved per step) x (number of steps). Pick a model that halves both and you beat a chip upgrade, for free.
This is arithmetic, not a benchmark — but it is the arithmetic that explains every "why is my Mac so slow" thread:
| Model | Params | Typical steps | Relative work per image (params x steps) |
|---|---|---|---|
| Z-Image-Turbo | 6B | 8 NFEs | 48 |
| SDXL base | 3B | 20-30 | 60-90 |
| FLUX.1 dev | 12B | 20-30 | 240-360 |
Parameter counts: Z-Image-Turbo model card ("6B parameters", "8 NFEs"); SDXL base 1.0 model card (3B); mflux README, which lists FLUX.1 as "legacy, 12B". Step counts are the values those projects ship as defaults. The right-hand column is a crude proportionality, not a timing.
FLUX dev is not four minutes per image because Macs are bad. It is slow because you asked for roughly five times the work of a turbo model, and then Apple's unified memory had to stream 12B parameters' worth of weights for every one of those steps. On a Mac, the step-efficient model is not a compromise — it is the correct default. We compare the two families in detail in SDXL vs FLUX locally.
What Each M4 Tier Actually Changes
Moving up the M4 ladder buys you memory headroom first and speed second — and on a Mac, memory headroom is what stops generation from falling off a cliff.
Apple Silicon has no separate VRAM pool. The GPU reads the same unified memory the OS and your apps use, which is why the VRAM figures in NVIDIA tutorials mislead Mac owners in both directions: a model that "needs 12GB VRAM" needs 12GB of your total RAM, but you also do not need a specific card to get it.
| Your Mac | Realistic image-gen posture |
|---|---|
| M4, 16GB (MacBook Air / base Pro) | Z-Image-Turbo class and SDXL are the targets. The 6B Z-Image-Turbo card explicitly claims it "fits comfortably within 16G VRAM consumer devices". Quit Chrome first — see Memory Pressure. |
| M4 Pro, 24-48GB | Comfortable SDXL at 1024x1024 with LoRAs, plus room to hold a quantized larger model without swapping. This is the tier where the workflow stops being a negotiation. |
| M4 Max, 64GB+ | Headroom for 12B-class models and long ComfyUI graphs held in memory at once. Faster than the Pro, but the honest gain over a well-fed M4 Pro is smaller than the price gap suggests. |
The one number we will not invent is the tok/s-equivalent for each tier on a given model. For the LLM side of the same question, our Apple Silicon buying guide and 16GB Mac model picks work through the unified-memory arithmetic tier by tier — but those are language models, and diffusion behaves differently enough that we are not going to launder one into the other.
Model Picks for a Mac
Start with Z-Image-Turbo. It is the only widely available model whose own card claims a 16GB consumer device and single-digit step counts at the same time.
Z-Image-Turbo (6B, Apache 2.0). The model card states 6B parameters, "8 NFEs (Number of Function Evaluations)", and that it "fits comfortably within 16G VRAM consumer devices". Apache 2.0 means no license drama for commercial output. On a Mac this combination is close to ideal: small enough to sit in unified memory beside macOS, and step-efficient enough that each image is a short job rather than a coffee break. Our Z-Image Turbo ComfyUI walkthrough covers the graph.
SDXL base (3B, CreativeML Open RAIL++-M). Still the most useful default when you need the enormous LoRA and fine-tune ecosystem that grew around it. It is also the only model with published Apple-measured timings, which makes it the right thing to benchmark your own machine with. Note the license is RAIL++-M, not Apache — read it before commercial use.
FLUX-family models. mflux lists FLUX.2 at 4B and 9B (Jan 2026) alongside legacy FLUX.1 at 12B, and Draw Things' own command-line example references a quantized FLUX.2 Klein 4B checkpoint (flux_2_klein_4b_q6p.ckpt), so the small FLUX.2 variants are a genuinely Mac-shaped option in a way 12B FLUX.1 dev never was. If you want the full FLUX.2 picture, see our FLUX.2 local setup guide.
What we are deliberately not recommending on a 16GB or 36GB Mac: 20B+ image models. mflux lists Qwen Image at 20B; it will load on a large Mac and it will be slow everywhere else. The broader ranking lives in best local image models compared.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
The Three Mac Stacks
Draw Things for "it just works", ComfyUI for node graphs, mflux if you want native MLX instead of PyTorch.
Draw Things — free, Metal, no terminal
Draw Things ships "for iPhone, iPad, and Mac". The Free Edition is genuinely free; the paid Draw Things+ tier at $8.99/month is for cloud compute, not for local generation, which matters because the usual "free app, paywalled features" suspicion does not apply to running models on your own machine. The community source is published under GPL-v3 in the drawthingsai/draw-things-community repository, which is a stronger transparency signal than most consumer AI apps offer. Its CLI example ships a quantized FLUX.2 Klein 4B checkpoint, so modern architectures are supported and not just SD 1.5.
We could not verify a complete, authoritative list of every architecture Draw Things currently supports — the marketing site does not publish one and the community README does not enumerate it. Check the in-app model catalogue before you plan a workflow around a specific checkpoint.
ComfyUI on MPS — the ecosystem, at the cost of an install
ComfyUI's README states you can install it "in Apple Mac silicon (M1, M2, M3 or M4) with any recent macOS version", and directs Mac users to install the latest PyTorch nightly via Apple's Accelerated PyTorch guide. That nightly requirement is the single most common reason Mac ComfyUI installs break, and it is still what the project's own README specifies:
# The nightly install command from Apple's Metal PyTorch page,
# which is what ComfyUI's manual-install instructions point you to
pip3 install --pre torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/nightly/cpu
# then, from the ComfyUI directory
pip install -r requirements.txt
python main.py
Worth knowing that this is a ComfyUI-specific requirement, not an Apple-wide one: Hugging Face's Diffusers documentation lists PyTorch 2.0 stable as recommended (1.13 as the minimum) for the mps device. If you are writing your own Diffusers script rather than running ComfyUI, you do not need nightly. Our ComfyUI complete guide covers the graph side once you are installed.
mflux — native MLX
mflux describes itself as running "the latest state-of-the-art generative image models locally on your Mac in native MLX", as "a line-by-line MLX port". Its published model list covers Z-Image (6B), FLUX.2 (4B & 9B), FLUX.1 (legacy, 12B), Qwen Image (20B), Krea 2 (12B), Ideogram 4 (9B), ERNIE-Image (8B), FIBO (8B), SeedVR2 and Depth Pro, with quantization support (its install example uses -q 8).
Is MLX faster than PyTorch MPS for diffusion? We do not know, and mflux does not claim it. Its README publishes no benchmark table, no timings and no memory figures — we checked specifically. Anyone telling you MLX is X% faster for image generation is quoting something other than the project. If it matters to your decision, run the benchmark recipe on both. For the LLM-side comparison, which is better documented, see MLX vs CUDA.
Memory Pressure Is What Actually Ruins Mac Generation
Hugging Face's Diffusers docs put it bluntly: "M1/M2 performance is very sensitive to memory pressure. When this occurs, the system automatically swaps if it needs to which significantly degrades performance." That single sentence explains most "my Mac got slower halfway through the batch" reports.
The fix Diffusers recommends is attention slicing, and its guidance is explicitly memory-size-based:
from diffusers import DiffusionPipeline
import torch
pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
torch_dtype=torch.float16,
variant="fp16",
use_safetensors=True,
).to("mps")
pipeline.enable_attention_slicing()
From the same doc, verbatim: attention slicing "usually improves performance by ~20% in computers without universal memory, but we've observed better performance in most Apple silicon computers unless you have 64GB of RAM or more." Read that carefully — on Apple Silicon under 64GB, slicing is a speed feature, not just a memory feature. Diffusers also recommends it "especially" if you generate at resolutions larger than 512x512, which is every SDXL job.
Two more documented gotchas from the same page: the MPS backend "does not support NDArray sizes greater than 2**32", and batching multiple prompts "can crash or fail to work reliably" — iterate instead of batching. If your ComfyUI queue dies at image four of eight, that is a candidate cause.
Practical checks before you blame the chip: quit the browser, watch the memory-pressure graph in Activity Monitor (or run memory_pressure in Terminal) while a generation is in flight, and re-run once nothing else is resident. If the graph goes yellow mid-generation, the number you just measured was a swapping number, not a hardware number. The same dynamic hits language models on Mac, which is why our Mac local AI setup guide leads with the same advice.
When It Errors Instead of Being Slow: The Three MPS Strings
Slow and broken are different problems. Everything above this line is about how long a Mac takes; if you got a traceback instead of an image, the error string tells you which of three unrelated failures you have — and only one of them is fixed by an environment variable. Match yours below, then take the fix from our dedicated ComfyUI on Mac: MPS errors and what fixes them page, which traces each string back to the PyTorch source line that emits it.
| The string you pasted into Google | What it actually means | Fixable with a flag? |
|---|---|---|
NotImplementedError: The operator 'aten::_int_mm' is not currently implemented for the MPS device. | An op in your graph has no Metal kernel at all. The operator name inside the quotes changes from workflow to workflow; the shape of the message does not. | Yes — PYTORCH_ENABLE_MPS_FALLBACK=1, but it has to be in the environment at launch, not exported afterwards |
TypeError: Cannot convert a MPS Tensor to float64 dtype as the MPS framework doesn't support float64. Please use float32 instead. | Metal has no float64 at all. A node asked for a double-precision tensor on the GPU, and it failed on tensor creation before any kernel dispatched. | No. No variable touches it — a file gets patched |
RuntimeError: MPS backend out of memory (MPS allocated: ..., other allocations: ..., max allowed: ...) | You collided with the MPS allocator's high watermark, which is a ratio of Metal's recommended working-set size — not your physical RAM. This is why it fires on a 64GB Mac with memory to spare. | Partly, via the watermark variables — and read the "other allocations" figure before you touch them |
The first two are quoted from PyTorch itself rather than from a forum: the fallback message is emitted by aten/src/ATen/mps/MPSFallback.mm, and the float64 one is the MPS_ERROR_DOUBLE_NOT_SUPPORTED macro in aten/src/ATen/mps/EmptyTensor.cpp. The watermark variables in the third row are documented in PyTorch's MPS environment variables reference.
Two more Mac-only failure modes are not in that table because they do not look like the three above. An fp8 checkpoint throws Trying to convert Float8_e4m3fn to the MPS backend but it does not have support for that dtype. — a different file, not a different flag. And the worst case throws nothing at all: the run reports success and the image comes out black or progressively corrupted. Both are worked through on the MPS errors page; we are deliberately not duplicating the fixes here, because a half-copy of a troubleshooting procedure is how people end up applying the wrong one.
One thing worth carrying back to this page: an out-of-memory error and the memory pressure problem above are not the same event. Pressure makes a run slow by swapping. The watermark makes it stop. If your generations are merely crawling, the fix is in Activity Monitor, not in an environment variable.
Benchmark Your Own Mac in Ten Minutes
Because nobody publishes M4 diffusion timings, the fastest route to a trustworthy number is to produce one. Use Apple's own benchmark configuration so your result is comparable to the M2 figures above:
- Fix the job. SDXL base, 1024x1024, 20 steps, classifier-free guidance on. Those are Apple's macOS benchmark settings — change any of them and you can no longer compare to the 37 s / 20 s figures.
- Warm up. Discard the first generation entirely. Model load, shader compilation and Metal warm-up all land on run one.
- Run five, take the median. Apple reports "the median latency value across 5 back-to-back end-to-end executions" for exactly this reason.
- Watch memory pressure the whole time (Activity Monitor, Memory tab). A yellow or red graph invalidates the run.
- Record what you changed. Attention slicing on or off, MPS or MLX or Core ML, and your macOS version. A seconds-per-image number without those four facts is not reproducible.
If you want a second data point to sanity-check against, repeat step 1 with Z-Image-Turbo at its default 8 NFEs. The gap between those two numbers on your own machine is the most useful thing this whole page can get you — it tells you, in seconds, what model choice is worth on your specific hardware.
What We Cannot Tell You
Being straight about the holes, because a page like this is easy to fake:
- We have not benchmarked M4, M4 Pro or M4 Max diffusion ourselves. Every timing on this page is Apple's, on M2-class hardware, and labelled as such.
- No verified MLX-vs-MPS speed comparison exists that we could find in either project's documentation. mflux publishes no benchmarks.
- Draw Things' full supported-model list is not published on its site or in the community README. Check the in-app catalogue.
- FLUX.2 Klein's parameter count and step defaults come from mflux's model table and Draw Things' example filename, not from a first-party model card we could read — the Hugging Face repository is gated.
- Peak memory figures per model on Mac are not published by any of the three stacks. Watch Activity Monitor; do not trust a number someone quotes you without a machine attached to it.
The gating dependency here is measurement, and we would rather ship a page that tells you how to measure than one that invents the measurement.
Sources
- Apple ml-stable-diffusion — first-party Apple Silicon latency benchmark tables (SD 2.1, SDXL; M2-class devices)
- ComfyUI README — Apple Mac silicon support statement (M1-M4) and PyTorch-nightly install instruction
- Diffusers MPS documentation — PyTorch version requirements, attention slicing guidance, memory-pressure and batching warnings
- PyTorch MPS environment variables —
PYTORCH_ENABLE_MPS_FALLBACKand the high/low watermark ratios behind the out-of-memory error - Z-Image-Turbo model card — 6B parameters, 8 NFEs, Apache 2.0, 16G consumer-device claim
- SDXL base 1.0 model card — 3B parameters, CreativeML Open RAIL++-M license
- mflux — MLX model list and quantization flags
- Draw Things and draw-things-community — platforms, free tier, $8.99/month cloud tier, GPL-v3 source
FAQ
Generating images locally? Take it further.
From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.
Go from one-off images to a real workflow
The Local Image Generation course covers ComfyUI, SDXL and FLUX properly — plus 24 more courses on running AI on your own hardware.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARRun FLUX.1 Locally in 2026: VRAM Needs + 5-Minute Setup
- AI-Toolkit LoRA Training: FLUX.2, Z-Image & Qwen-Image
- Best GPU for Local AI Image Generation (2026): Ranked
- Best Local AI Image Models 2026: FLUX vs SDXL vs Qwen
- blog/flux-vram-requirements-by-gpu
- Chroma Local Guide: The Apache-2.0 Uncensored FLUX Model
- ComfyUI Black Image Fix: NaN, VAE and fp8 by Model
- ComfyUI FLUX Workflow (2026): JSON Nodes Explained
- ComfyUI IMPORT FAILED: Find the Real Error Fast
- ComfyUI LoRA Not Working: Key Not Loaded Fixes
Comments (0)
No comments yet. Be the first to share your thoughts!