Can Your GPU Run FLUX.2, Qwen-Image or Z-Image?
At 1024x1024, an 8GB card runs Z-Image (GGUF Q4_K_M, ~6.9 GB peak) and FLUX.2 [klein] 4B at FP8 (~6.2 GB, set by the encode phase); 12GB adds klein 9B at FP8 (~11.2 GB); Qwen-Image FP8 peaks at ~22.9 GB on a 24GB card, and FLUX.2 [dev] GGUF Q4 denoises in ~22.1 GB but its FP8 encoder needs ~24.6 GB, so 24GB only works with the encoder pushed to system RAM. Time per image on a 4090-class GPU runs from ~1 second (klein 4B, 4 steps) to ~56 seconds (FLUX.2 [dev], 28 steps). Every weight figure below is a published file size, not a rule of thumb — and the encoder-offload toggle is the switch that decides whether an 8GB card works at all.
No — needs 14.2 GB, you have 12 GB
2.2 GB short · 28 steps x 2 passes at 0.39 s per pass
The transformer itself does not fit (14.2 GB of denoise-phase memory). ComfyUI will stream weights from system RAM every step. It usually still produces an image — it just stops being interactive. Drop to a smaller build or a lower resolution.
How the two numbers differ in confidence
The VRAM figure is arithmetic on published file sizes; the time figure is a one-constant fit. Seconds per pass here is 0.0625 x parameters(B) x megapixels on an RTX 4090-class card, fitted to three of our own timings and accurate to roughly a factor of two. Generate one image, divide your real time by this estimate, and set the slowdown slider to that number — after which every other row on this page becomes right for your machine.
Licence: Apache-2.0. real CFG — two forward passes per step. Weight data: Tongyi-MAI/Z-Image (6B, Apache 2.0) + Comfy-Org/z_image and unsloth/Z-Image-GGUF file listings.
The Formula, and Why the Answer Moves
A modern image model is four separate memory costs that peak at different moments, which is why one number never fits. The transformer, the text encoder, the VAE and the activation working set are loaded in a sequence, and ComfyUI evicts the encoder once your prompt has been turned into embeddings.
Latent tokens (VAE downsamples 8x, DiT patchifies 2x2):
tokens = (width / 16) x (height / 16)Peak with the encoder offloaded (the ComfyUI default):
max( encoder + overhead , transformer + vae + activations + overhead )Peak with everything resident:
transformer + encoder + vae + activations + overheadTime:
seconds = steps x cfg_passes x 0.0625 x parameters(B) x megapixels x your_slowdownThe max() in the offload case is the whole reason 8GB cards work. It also explains a result that surprises people on a 24GB card: FLUX.2 [dev] as a GGUF Q4_K_S needs about 22.1 GB to denoise, which fits — but its FP8 Mistral Small 24B text encoder is roughly 24 GB by itself, so the encode phase is the binding constraint, not the diffusion phase. That is why FLUX.2 [dev] workflows commonly push the encoder to system RAM.
The arithmetic reproduces the vendors' own headline claims, which is the check that it is real rather than curve-fitted. Tongyi says Z-Image “fits comfortably within 16G VRAM”; the components give 12.31 + 0.34 + 0.9 + 0.6 = 14.2 GB. The Qwen-Image repo totals 57.7 GB; the BF16 transformer (~40.9 GB) plus the FP16 Qwen2.5-VL-7B encoder (~16 GB) plus the VAE lands there. Black Forest Labs' ~13 GB figure for a BF16 klein 4B is the 8.04 GB transformer plus the 5.63 GB FP8 Qwen3-4B encoder held together.
The activation term is an allowance, not a measurement. We use 0.6 GB plus 0.05 GB per billion parameters, per megapixel — about 0.9 GB at 1 MP for a 6B model and 2.2 GB for a 32B one. Real peak working memory depends on your attention backend; a non-memory-efficient path grows quadratically with token count and will exceed this at 2048px. Treat a result within ~1 GB of your card as tight.
Weights, From the Published Files
These are file sizes read off Hugging Face, not parameters multiplied by a nominal bit width. Where a size is marked derived, we multiplied the parameter count by the effective bytes-per-parameter computed from a published file of the same dtype (BF16 = 2.01 bytes, from FLUX.2 [dev]'s 64.4 GB / 32B).
| Model | Params | Full precision | Smallest build | Text encoder | Licence |
|---|---|---|---|---|---|
| FLUX.2 [dev] | 32B | 64.4 GB BF16 | ~19 GB GGUF Q4_K_S | Mistral Small 24B (24 GB FP8) | Non-commercial |
| FLUX.2 [klein] 9B | 9B | ~18.1 GB BF16 (derived) | ~9.2 GB FP8 (derived) | Qwen3-8B (8 GB FP8) | Non-commercial |
| FLUX.2 [klein] 4B | 4B | ~8.04 GB BF16 (derived) | ~4.1 GB FP8 (derived) | Qwen3-4B (5.63 GB FP8) | Apache-2.0 |
| Qwen-Image | 20B | ~40.9 GB BF16 | ~8 GB Nunchaku 4-bit | Qwen2.5-VL-7B (9 GB FP8) | Apache-2.0 |
| Qwen-Image-Edit-2509 | 20B | ~40.9 GB BF16 | ~12.5 GB GGUF Q4_K_S | Qwen2.5-VL-7B (9 GB FP8) | Apache-2.0 |
| Z-Image / Z-Image-Turbo | 6B | 12.31 GB BF16 | 4.01 GB GGUF Q2_K | Qwen3-4B (5.63 GB FP8) | Apache-2.0 |
Sizes from the black-forest-labs/FLUX.2-dev, FLUX.2-klein-4B, Qwen/Qwen-Image, Qwen/Qwen-Image-Edit-2509, Tongyi-MAI/Z-Image, Comfy-Org/z_image and unsloth/Z-Image-GGUF repositories, read 18 August 2026. The licence column matters as much as the size column for client work: only FLUX.2 [klein] 4B is permissively licensed in the FLUX.2 family, while every Qwen-Image and Z-Image build here is Apache-2.0. Our local image model comparison weighs that against output quality.
What Each Card Tier Actually Runs
At 1024x1024 with an FP8 encoder and ComfyUI's default encoder offload, the tiers land like this. These are the calculator's own outputs, reproduced so you can see the shape before touching a control.
| Your VRAM | Best quality that fits | Peak | ~Time (4090-class) |
|---|---|---|---|
| 8 GB | Z-Image-Turbo GGUF Q4_K_M, 8 steps | ~6.9 GB | ~3 s |
| 8 GB | FLUX.2 [klein] 4B FP8, 4 steps | ~6.2 GB (encode phase) | ~1 s |
| 12 GB | FLUX.2 [klein] 4B BF16, 4 steps | ~9.8 GB | ~1 s |
| 12 GB | FLUX.2 [klein] 9B FP8, 4 steps | ~11.2 GB | ~2 s |
| 16 GB | Z-Image base BF16, 28 steps at CFG 4 | ~14.2 GB | ~21 s |
| 16 GB | Qwen-Image GGUF Q4_K_S, 20 steps | ~15.0 GB | ~25 s |
| 24 GB | Qwen-Image FP8, 20 steps | ~22.9 GB | ~25 s |
| 24 GB | FLUX.2 [dev] GGUF Q4_K_S, 28 steps — only with the encoder on the CPU | ~22.1 GB denoise / ~24.6 GB encode | ~56 s |
| 32 GB | FLUX.2 [dev] GGUF Q4_K_S, 28 steps — encoder now fits too | ~24.6 GB | ~56 s |
| 48 GB | FLUX.2 [dev] FP8, 28 steps | ~35.1 GB | ~56 s |
The FLUX.2 [dev] row on 24GB is the honest one: the transformer fits at ~22.1 GB and the FP8 Mistral encoder does not leave room, so a 24GB card needs the encoder on the CPU. Per-card model picks live on our 12GB, 16GB and 24GB FLUX pages.
Why Steps Beat Parameters on Time
A 4B distilled model at 4 steps is roughly 20x faster than a 6B undistilled one at 28 steps with CFG — the size difference is irrelevant next to the pass count. Two things multiply here and both are easy to miss.
- Step distillation. FLUX.2 [klein] ships with
num_inference_steps=4in its own card example. Z-Image-Turbo runs 8 NFEs. Both trade output diversity for a 5-10x cut in compute. - Classifier-free guidance doubles every step. Real CFG above 1.0 means a conditional and an unconditional forward pass. Z-Image base at 28 steps and CFG 4 is 56 passes. Qwen-Image-Edit-2509's card example uses 40 steps at
true_cfg_scale 4.0— 80 passes, which is why editing feels so much slower than generating. - Resolution is linear on time and near-linear on memory. 2048x2048 is four times the latent tokens of 1024x1024, so four times the compute per pass.
The practical workflow most people converge on: prototype prompts on a distilled model, then re-render the keepers on the undistilled one. Same encoder, same VAE, one swapped checkpoint file — the ComfyUI guide covers building that as a single graph with a bypass toggle.
Where the Time Estimate Comes From
One constant, sanity-checked against a handful of secondhand timings, and it is accurate to roughly a factor of two. None of these are our numbers. We do not own the cards in the table below — the anchors are times reported elsewhere for these models, which we have not verified. We publish the fit rather than a benchmark grid because inventing cells would be worse than admitting the gap.
| Anchor | Reported elsewhere | Formula says | Provenance |
|---|---|---|---|
| Z-Image-Turbo, 8 steps, 1024px, RTX 4090 | 2-3 s | 3.0 s | secondhand, unverified by us |
| FLUX.2 [klein] 4B, 4 steps, capable GPU | ~1 s | 1.0 s | secondhand, unverified by us |
| FLUX.2 [dev] GGUF Q4, 1 MP, RTX 4090 | tens of seconds | 56 s at 28 steps | secondhand and imprecise |
| Qwen-Image Nunchaku 4-step, RTX 3060 12GB | 8-15 s | ~5 s at 1x, ~20 s at 4x slowdown | secondhand; card factor unknown |
The slowdown slider exists because we deliberately did not publish per-card multipliers. Generate one image, divide your real time by the estimate, set the slider to that ratio, and every other row on this page becomes correct for your machine. That takes about forty seconds and beats any table we could have guessed at.
What This Calculator Does Not Know
- Quantised is not always faster. A GGUF build fits in less VRAM but dequantises on the fly, and on some backends it is slower per step than an FP8 build that also fits. The time estimate here is driven by parameters and pixels, not by the quant format.
- Apple Silicon breaks the model. On unified memory there is no VRAM/RAM split to optimise, so the offload toggle is meaningless and throughput is governed by memory bandwidth rather than by a discrete GPU's compute. Treat the VRAM number as a total-memory floor and ignore the time estimate.
- LoRAs, ControlNets and upscalers are extra. Each adds weights and, for tiled upscaling, a second full pass. None of it is modelled here.
- Some sizes are derived, not published. The FLUX.2 [klein] rows and the FP8 variants are parameter counts times a bytes-per-parameter figure computed from a published file — they are marked derived in the picker for that reason.
- Model cards move. Variant names and step recommendations in this family change with each release; re-check the card before you download 40 GB.
Once You Know What Fits
The next step is the download list for the variant you just picked. For the FLUX.2 family, the FLUX.2 local setup guide has the exact file names and ComfyUI folders; if the answer above was “no”, the low-VRAM FLUX guide covers the offload flags that buy you another tier. For Z-Image, start with Z-Image Turbo in ComfyUI; for the 20B text-rendering specialist, the Qwen-Image local guide. If you are still choosing hardware, best GPU for image generation ranks cards by image throughput rather than LLM specs, and the LLM VRAM calculator is the text-model equivalent of this page.
Frequently Asked Questions
What are the FLUX.2 VRAM requirements?
They differ by an order of magnitude across one family, which is why a single "FLUX.2 VRAM" number is always wrong. FLUX.2 [dev] is a 32B transformer and its BF16 file is 64.4 GB; the FP8 build is about 32 GB and a GGUF Q4_K_S about 19 GB, so 24GB is the entry point and even then the 24 GB FP8 Mistral Small 24B text encoder is tighter than the transformer. FLUX.2 [klein] 9B at FP8 needs roughly 11 GB during denoising, so a 12GB card runs it. FLUX.2 [klein] 4B is the one most people can actually run: about 9.8 GB peak at BF16 and about 6.2 GB at FP8 with the encoder offloaded — at FP8 the encode phase, not the transformer, sets the peak — and it is the only FLUX.2 weight under Apache-2.0. Note that BFL quotes higher headline figures on its own cards (~13GB for klein 4B, ~29GB for klein 9B); those assume reference precision with the text encoder resident, which is exactly what the offload toggle here removes.
Can an 8GB GPU run Z-Image?
Yes, on a quantised build. The BF16 diffusion checkpoint is a 12.31 GB file, which lands at roughly 14 GB of peak memory at 1024x1024 and wants a 16GB card — matching Tongyi's own "fits comfortably within 16G VRAM" statement. The GGUF Q4_K_M build is a 5.07 GB file, which peaks around 6.9 GB once ComfyUI has evicted the text encoder — comfortable on 8GB. The int8 build at 6.20 GB peaks right on 8.0 GB, so treat that one as tight rather than safe. The catch on 8GB is the encoder, not the model: the FP8 Qwen3-4B encoder is 5.63 GB on its own, so it has to run and then unload before the transformer loads. That sequencing is the ComfyUI default, which is why this works at all.
How long does it take to generate an image locally?
At 1024x1024 on an RTX 4090-class card, the honest spread runs from about one second to about a minute, and step count matters more than model size. FLUX.2 [klein] 4B is step-distilled to 4 steps: roughly a second. Z-Image-Turbo at 8 steps: 2-3 seconds. Z-Image base runs 28-50 steps with real CFG, and CFG above 1.0 costs two forward passes per step, so ~28 steps is really 56 passes — around 21 seconds. Qwen-Image at 20 steps is roughly 25 seconds and Qwen-Image-Edit-2509 at its documented 40 steps with true CFG 4.0 is around 100. FLUX.2 [dev] at 28 steps is around a minute. Slower cards scale roughly linearly; a model that does not fit is the one case where the estimate collapses.
Does resolution change how much VRAM I need?
Yes, but less than people expect, and it changes generation time much more. The weights are a fixed cost — a 12.31 GB checkpoint is 12.31 GB at any resolution. What grows is the activation working set, which scales with the number of latent tokens: the VAE downsamples by 8 and the DiT patchifies 2x2, so a 1024x1024 image is 4,096 tokens and a 2048x2048 image is 16,384. Our allowance for that working set is about 0.9 GB at 1 MP for a 6B model and 2.2 GB for a 32B model, scaled linearly with pixel count. Time, by contrast, scales directly with megapixels: 4 MP costs four times the compute of 1 MP at the same step count.
Why does the calculator say it fits when ComfyUI still runs out of memory?
Three usual reasons, and none of them are the weights. First, the attention backend: with plain PyTorch SDPA and no memory-efficient path, attention memory grows quadratically with token count, so 2048x2048 can blow past a linear allowance. Second, a pinned text encoder — if your workflow keeps the encoder resident instead of letting ComfyUI evict it, add the full encoder size back. Third, other processes: a browser with hardware acceleration and a desktop compositor can hold a gigabyte before you start. Treat a result inside about 1 GB of your card as "tight", not "fits", and drop one build tier.
Sources
- black-forest-labs/FLUX.2-dev — “32 billion parameter rectified flow transformer”, FLUX Non-Commercial License; file listing gives flux2-dev.safetensors 64.4 GB and ae.safetensors 336 MB
- black-forest-labs/FLUX.2-klein-4B — 4B rectified flow transformer,
apache-2.0,num_inference_steps=4 - Qwen/Qwen-Image (20B params, Apache 2.0, repo total 57.7 GB) and Qwen-Image-Edit-2509 (
num_inference_steps 40,true_cfg_scale 4.0) - Tongyi-MAI/Z-Image-Turbo — 6B, Apache 2.0, “8 NFEs”, “fits comfortably within 16G VRAM consumer devices”; Tongyi-MAI/Z-Image for the undistilled base
- Comfy-Org/z_image and unsloth/Z-Image-GGUF file listings — z_image_bf16 12.31 GB, int8 6.20 GB, Q8_0 7.22 GB, Q4_K_M 5.07 GB, Q2_K 4.01 GB, qwen_3_4b encoder 8.04 GB (FP8 5.63 GB), ae 0.34 GB
- Our own timings, reported in the Z-Image Turbo, FLUX.2, FLUX VRAM and Qwen-Image guides — approximate single-machine observations, not controlled benchmarks
All repositories read 18 August 2026.
Free to use — just keep the attribution link. Works on any site.
<iframe src="https://localaimaster.com/embed/image-model-vram-calculator" width="100%" height="560" style="border:1px solid var(--line);border-radius:12px;max-width:680px" title="Image Model VRAM Calculator — Local AI Master" loading="lazy"></iframe>
<p style="font:13px/1.5 system-ui,sans-serif;max-width:680px;margin:6px 0 0"><a href="https://localaimaster.com/tools/image-model-vram-calculator">Image Model VRAM Calculator</a> by <a href="https://localaimaster.com">Local AI Master</a></p>Know what to actually run on it
All 561 chapters — running local models, RAG, agents, fine-tuning — plus the Python Lab and every course added later.
Ready to Go Beyond Tutorials?
25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.
Was this helpful?
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Written by the Local AI Master Team
The team behind Local AI Master
We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.
Related Guides
Continue your local AI journey with these comprehensive guides
FLUX.2 Local Setup Guide
The download list and ComfyUI workflow for dev, klein 9B and klein 4B.
Z-Image Turbo in ComfyUI
The 8-step model that made a single GPU feel interactive.
Qwen-Image Local Guide
The 20B text-rendering specialist, from BF16 down to 8GB.
Run FLUX on a Low-VRAM GPU
Offload flags and quant choices when it does not fit.
Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide
No spam. Unsubscribe with one click.