ComfyUI Black Image Fix: NaN, VAE and fp8 by Model
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Generating images locally? Take it further. From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.
A pure black output with no error message means your latent went to NaN before the VAE decoded it, and on most machines the single fastest fix is to restart ComfyUI with --fp32-vae. That is not folklore: ComfyUI's own argument parser describes the opposite flag, --fp16-vae, as "Run the VAE in fp16, might cause black images." If --fp32-vae does not fix it, the cause is specific to the model, the precision and the backend you are running: an fp16-overflowing SDXL VAE, an fp8 or int-quantised checkpoint without the matching kernels, a bf16 weight path on a brand-new 2026 checkpoint, a warm model handed back by the memory manager, or a checkpoint that comes out black at every precision you try. Grey or brown output is a different bug entirely — that is the wrong VAE, not a NaN.
The reason this failure is so hard to search is that nothing fails. The queue completes, the progress bar fills, the node graph goes green, and the SaveImage node writes a valid PNG that happens to be 100% #000000. There is no traceback to paste. This page gives you the thread to pull: a two-node test that proves whether it is NaN, then the nine causes ordered by how often each one turns out to be the real one, then a model / precision / backend lookup table you can scan straight to your own row — because the fix for SDXL actively makes the FLUX case worse, and the 2026 model releases inverted the rule that "less quantisation is safer".
Flag names below were read from comfy/cli_args.py on master at the time of writing (ComfyUI v0.33.1, released 13 August 2026 — that file has changed a lot across the v0.3x series, so check yours with python main.py --help). Bug reports are linked to their issue numbers so you can check whether yours has been fixed since.
A note on how the newest reports are cited. Four of the 2026 cases below are quoted by their tracker title rather than by a number, because we could not confirm a stable issue number for them. Where that is the case we say so and link a tracker search for the title instead of a specific issue — a search that finds the live thread is more useful to you than a number we are not certain of.
What Error Message Should I Be Looking For?
The black image itself produces no error. But in a huge share of these reports there is one line in the console, and it is the giveaway:
RuntimeWarning: invalid value encountered in cast
img = Image.fromarray(np.clip(i, 0, 255).astype(np.uint8))
That warning appears in issue #4673 (FLUX dev fp8 outputting pure black on a 4090) and again in issue #10681 (black images on an M3 Ultra Mac Studio). It means NumPy was handed NaN and cast it to an integer. NaN clips to zero. Zero is black.
If you see that line, stop reading about prompts, samplers, CFG and negative embeddings. Your pipeline produced not-a-number and the rest of this page applies. If you do not see it, scroll up in the console anyway — it is a warning, not an error, so it prints once and gets buried by the next queue item.
Coming from Automatic1111 or Forge instead? The equivalent knob there is --no-half-vae, and the same underlying overflow is what it works around.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
How Do I Confirm It Is a NaN and Not a Bad Prompt?
Two nodes settle it in under a minute, and you should do this before changing anything.
- Bypass the VAE. Wire your KSampler's LATENT output into a PreviewImage through a fresh VAEDecode, but first drop the sampler to 1 step and set CFG to 1.0. A 1-step, CFG-1 generation should produce a blurry, ugly, but coloured image. If 1 step is coloured and 20 steps is black, the NaN is being produced during sampling — cause 3, 5, 6 or 7 below. If even 1 step is black, the decode is the problem — cause 1 or 2.
- Move the decode to CPU. Restart with
--cpu-vae(the help text is simply "Run the VAE on the CPU."). CPU tensors are fp32 by default. If the image comes back correctly on the CPU VAE, you have proven a GPU-side precision overflow and--fp32-vaeis your permanent fix. If it is still black on the CPU, the latent arriving at the VAE was already NaN and the VAE is innocent.
That second test is the most useful diagnostic on this page, because it cleanly splits "the decode broke" from "the sample broke" — and every fix below sits on one side or the other of that line.
What Causes a Black Image in ComfyUI? The Nine, Ranked
| # | Cause | Tell | Fix |
|---|---|---|---|
| 1 | SDXL-family VAE overflows in fp16 | Black only after decode; CPU VAE works | --fp32-vae, or load sdxl-vae-fp16-fix |
| 2 | Wrong or missing VAE for the checkpoint | Output is grey/brown/washed, not pure black | Load the VAE the model was trained with |
| 3 | fp8 / int-quantised weights without matching kernels | Black on one GPU, fine on another, same file | Use the fp16/bf16 build, or a GGUF quant |
| 4 | Apple Silicon MPS precision | Black on a Mac, same graph fine on CUDA | --force-fp32, or --cpu-vae |
| 5 | Warm model reuse under dynamic VRAM | First run good, every later run black | Restart without dynamic VRAM; unload between runs |
| 6 | Hook/LoRA weights lost after offload | Black only when a hook node is in the graph | Use LoraLoaderModelOnly; disable async offload |
| 7 | bf16 path on a brand-new checkpoint | Black on bf16 file, fine on the quantised repack | Use the vendor's quantised build for now |
| 8 | AMD ROCm allocator returning zeros | Black and white frames, after a KFD fault in dmesg | Drop expandable_segments:True |
| 9 | The checkpoint is black at every precision | fp16, bf16 and fp32 all black; a sibling build works | Model-side bug: use the sibling build, report yours |
Work down that list in order. Causes 1 and 2 account for the overwhelming majority of what people are actually hitting; 5 through 9 are current-release bugs that will churn — and cause 9 is the one that did not exist before the late-2026 model wave, when several vendors shipped a fast variant that works and a full variant that does not.
Cause 1: The SDXL fp16 VAE Overflow (Start Here)
This is the classic, and it is a property of the original SDXL VAE weights, not of your install. The sdxl-vae-fp16-fix model card states the problem in one sentence: "SDXL-VAE generates NaNs in fp16 because the internal activation values are too big." fp16's maximum representable value is 65,504; the decoder's internal activations exceed it, the tensor goes to infinity, infinity minus infinity is NaN, and everything downstream is NaN.
Two fixes, and they are not equivalent:
Fix A — run the VAE in fp32 (universal, costs VRAM):
python main.py --fp32-vae
This works for every model, not just SDXL, because it removes the overflow condition entirely. The cost is real: the decode step holds fp32 activations, which on a tight 8GB card can push you from a black image into an out-of-memory error instead. If that happens, see our guide to ComfyUI's memory manager and dynamic VRAM for the offloading knobs — trading one failure for another is progress, but only just.
Fix B — use the rescaled VAE (cheaper, SDXL only): download sdxl.vae.safetensors from madebyollin/sdxl-vae-fp16-fix into ComfyUI/models/vae/, add a VAELoader node, and wire it into your VAEDecode instead of the checkpoint's built-in VAE. The author fine-tuned the VAE to produce the same final output with smaller internal activations by "scaling down weights and biases within the network," so it simply never reaches the overflow. The model card is honest about the tradeoff: "There are slight discrepancies between the output of SDXL-VAE-FP16-Fix and SDXL-VAE, but the decoded images should be close enough for most purposes."
Use Fix B if you generate SDXL all day and want your VRAM back. Use Fix A if you switch between model families, because Fix B does nothing for FLUX, Qwen-Image or a video model.
One warning that applies to both: do not "fix" this by adding --fp16-vae. People find that flag while searching and try it because it looks like a precision setting. It is the setting that causes the bug, which is why its own help text says so.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
Cause 2: Grey, Brown or Washed-Out Is a VAE Mismatch, Not a NaN
If your output is a flat grey or muddy brown rather than true black, stop reading about NaN — you decoded the latent with the wrong VAE. This is the second-most-common report and it gets misfiled under "black image" constantly, because at a glance a dark grey 512×512 square looks black in a file browser thumbnail.
Every model family uses its own latent space. SD 1.5 latents decoded with an SDXL VAE, or FLUX latents decoded with an SD VAE, produce structured garbage — often a uniform mid-tone with faint colour blocking — because the decoder is interpreting the numbers correctly according to a different mapping. The check: open the PNG in an editor and read the pixel values. Exactly 0,0,0 across the whole frame is NaN. Anything else — 40,38,35 or a smeary brown — is a VAE mismatch.
The fix is to load the matching VAE explicitly rather than relying on whatever is baked into the checkpoint:
- SD 1.5: the checkpoint's built-in VAE is usually fine;
vae-ft-mse-840000-ema-prunedis the common upgrade. - SDXL and its descendants:
sdxl_vae.safetensors, or the fp16-fix version above. - FLUX: the FLUX autoencoder (
ae.safetensors) — nothing else will decode a FLUX latent. Our FLUX VRAM requirements guide lists the support files alongside each diffusion model, and the FLUX.2 setup guide covers the current generation's file layout. - Qwen-Image / Z-Image and the video models: each ships its own VAE file. Do not reuse one across families. The Z-Image Turbo ComfyUI guide walks through the correct loader wiring for that family.
A related trap: some checkpoints are distributed without a baked VAE. ComfyUI will happily run the graph and hand you the mismatch rather than erroring, so a missing VAE and a wrong VAE look identical from the outside.
Cause 3: fp8 and Int-Quantised Checkpoints Without the Kernels
The symptom that identifies this one: the exact same file produces a good image on one GPU and a black image on another. Quantised weights are only half the story — the other half is a kernel that can multiply them, and that kernel is compiled for specific architectures.
The history in ComfyUI's own tracker is instructive:
- #4572, "SDXL generate black images with new --fast arg" (August 2024) was fixed by PR #5928, titled "Fix SDXL generating black images when using --fast (fp8_ops)". The
--fastpath swapped in fp8 matmuls that SDXL's numerics could not survive. - #4673, "The flux dev model outputs pure black on 4090", reports the fp8 e4m3fn FLUX dev build going black on an RTX 4090 while FLUX Schnell and other models are fine on the same machine, with the "invalid value encountered in cast" warning in the log. It is still open.
- #9046, "Black images in XL and Nunchuka(Kontext Dev) caused by Ksampler outputting NaNs. Nothing works.", ties the same failure to an INT4 inference backend on an RTX 3080 Ti. It is still open.
What to do about it, in order:
- Swap the file, not the flags. If the fp8 build is black, download the fp16 or bf16 build of the same model and confirm it works. That single test tells you whether you are looking at a quantisation problem or something else, and it costs you a download rather than an afternoon.
- If fp16 works and fp8 does not, use a GGUF quant instead. GGUF quants dequantise to a dtype your card definitely supports rather than relying on native fp8 tensor cores. You give up some speed and get correctness.
- Do not add
--fast. If you have it in your launch script from an old tutorial, remove it and retest — it is precisely the sort of aggressive-op flag that produced #4572. - Check the arch, honestly. Native fp8 matmul is a Blackwell/Ada-class capability; on older cards ComfyUI stores fp8 weights and upcasts them, which is a different code path with different numerics. We do not own one of every GPU generation and we are not going to publish a card-by-card kernel matrix we cannot verify — the reliable test is the fp16-vs-fp8 swap above, run on your GPU. If the quantised build is failing because it will not fit rather than because it will not compute, our low-VRAM FLUX guide covers the offload options that keep you on a dtype your card actually supports.
The explicit dtype flags exist if you want to force the issue during debugging: --fp8_e4m3fn-unet, --fp8_e5m2-unet, --fp16-unet, --bf16-unet, --fp32-unet, plus the text-encoder equivalents --fp8_e4m3fn-text-enc, --fp16-text-enc and --fp32-text-enc. Forcing --fp16-unet on a model that expects bf16 is itself a way to manufacture a black image, so change one thing at a time.
Cause 4: Why Do I Only Get Black Images on a Mac?
On a Mac, try --force-fp32 before anything else, then --fp32-vae, then --cpu-vae. ComfyUI's own log line when it engages the first of those reads "Forcing FP32, if this improves things please report it," which tells you how the maintainers regard MPS numerics. Each step costs memory and speed, and on unified memory that cost is felt immediately.
Two things are worth knowing before you start swapping flags. First, issue #10681 — black images on a Mac Studio (M3 Ultra, 256GB) with a WanVAE / Qwen-Image graph while the same workflow is correct on an M3 Max MacBook Pro — reproduces in a separate app too, so it is an MPS-level problem rather than a ComfyUI bug. Second, the 2026 video models added a distinctly different Mac symptom: an open tracker report describes random all-NaN attention on MPS producing black videos with bf16 LTX-2.x models, and proposes two fixes in-thread (search the tracker — this is one of the reports we are quoting by title, not by number). The report's own word is "random", so intermittency is the signature — a graph that renders correctly on one run and black on the next, with nothing changed, is this and not a precision setting you got wrong. If you are on that family, our LTX-2 local setup guide has the working file layout to compare against.
Everything else Mac-specific lives on its own page. Rather than duplicate it here, see ComfyUI on Mac: MPS errors and what fixes them for the backend errors, the fallback behaviour and the attention-implementation flags. (Historic note for searchers: #15315, the official MiniMax H3 text-to-video workflow producing black video and NaN audio on an M4 Max, has been closed.)
Cause 5: Warm Models Under Dynamic VRAM
Tell: the first generation after a fresh model load is perfect, and every subsequent generation from the same queue is black. That pattern is unmistakable once you know to look for it, and it is a live regression rather than a configuration mistake.
Issue #15452 — "Dynamic VRAM: reused (warm) model produces NaN/black output on VAE decode, fresh load does not", opened 9 August 2026 and open at the time of writing — reports that when prompts are queued back to back against the Boogu T2I Turbo official template, "every subsequent generation that reuses the already-resident Boogu model renders fully black", with the same "invalid value encountered in cast" warning. The reporter's environment was ComfyUI v0.31.0, an RX 7600 XT on ROCm 7.14, with dynamic VRAM enabled and the VAE already running in fp32 — note that fp32 VAE was already on, which is what makes this a distinct cause rather than cause 1 in disguise. Their evidence is a missing log line: runs that printed "Requested to load Boogu" succeeded, and runs that skipped straight to the dynamic-VRAM prepare step failed.
Workarounds, all unsatisfying: restart without the dynamic-VRAM flag; or unload models between generations instead of queueing prompts back to back, which throws away most of what dynamic VRAM is for. Check whether the issue has been closed before you rearrange your workflow around it.
If you arrived here from an out-of-memory error after an update rather than a black image, that is the same subsystem behaving differently and a different diagnosis — the memory-manager section of our complete ComfyUI guide is the better starting point.
Cause 6: Hooks and LoRAs That Vanish After Offload
Tell: the graph is black only when a hook node is present, and identical without it. Issue #15531, "BF16 WeightHooks can produce black output after model offload" (opened 12 August 2026, open), records a mapped WeightHook producing an all-black image on a BF16 Krea 2 checkpoint even though the prompt completes successfully.
The decisive detail for diagnosis: with DynamicVRAM enabled, the same LoRA works through LoraLoaderModelOnly and, through CreateHookLoraModelOnly → SetClipHooks, completes but yields what the reporter calls "black/non-finite output". Their diagnosis is that hook writeback after model preparation and offload has to preserve a valid storage and transfer lifecycle, and the candidate fix is to keep plain-weight hook transfers synchronous rather than pinned-async.
So: if your graph uses hook nodes, rebuild that section with the plain LoRA loader and retest. If that fixes it, you have confirmed the cause and can decide whether you need hooks badly enough to also disable async offload (--async-offload controls the number of offload streams and is on by default on NVIDIA).
Cause 7: bf16 Paths on the Newest Checkpoints
Tell: the vendor's quantised repack works and the plain bf16 file of the same weights does not. This is the 2026-specific version of the problem, and it shows up within days of each new model release.
Issue #15563 (opened 13 August 2026, open) is the cleanest example on record. The reporter tested three files of the same MiniMax H3 weights on ComfyUI v0.32.0, Windows 11, an RTX PRO 6000 Blackwell 96GB, torch 2.12.0+cu130:
| Checkpoint | Size | Result |
|---|---|---|
minimax_h3_ref2va_bf16.safetensors | 61.7 GB | 100% black, every frame |
minimax_h3_ref2va_pruned_bf16.safetensors | 37.5 GB | 100% black, every frame |
minimax_h3_ref2va_int8_convrot.safetensors | — | Correct output |
Their conclusion, after ruling out file corruption, attention backends, VAE precision and memory pressure: "the only path that differs between the working and failing runs is runtime op dispatch: int8_convrot layers execute through the comfy_kitchen quantized kernels, while plain bf16 layers execute through standard ops."
Note what this inverts. Everywhere else on this page, quantisation is the suspect and full precision is the safe fallback. Here it is the reverse — the quantised build is the one that works. The practical rule for any checkpoint released in the last few weeks: if one precision variant is black, try the others before you conclude anything about your hardware. Vendors ship several builds precisely because the fast paths are new code.
But do not stop at the transformer. The int8 repack that fixes the sampler has its own reported failure at the other end of the graph: a separate open report, titled in the tracker as the MiniMax H3 INT8 ConvRot VAE producing black video with NaN/Inf output, puts the non-finite values at the VAE stage rather than during sampling (quoted by title — search the tracker for the live thread). The practical consequence is that on this family the two halves of the pipeline want different things, so test them separately: run the --cpu-vae check from the top of this page, and if the quantised transformer plus a higher-precision VAE decode gives you a picture, keep that pairing. It is the same split the diagnostic at the top exists to expose, just on a model where both sides have live bugs at once.
Two adjacent reports from the same wave, worth knowing about so you do not misfile your own symptom: #15314 has MiniMax H3 producing pure noise (not black) on an RX 7900 XTX across every quantisation and backend combination, and #15617 has an INT8 ConvRot build hard-rebooting a Windows machine on an RTX 3080. Noise is not NaN, and a reboot is not a numerical failure.
Cause 8: AMD, ROCm and Silent Zeros
On an AMD card, check dmesg before you touch ComfyUI. There is a ROCm-level failure mode that produces exactly the symptom this page is about, and no amount of VAE precision will fix it.
ROCm issue #6603 (RX 9070 XT, gfx1201, ROCm 7.14.0, a regression from 7.2.4) reports that after a KFD userptr restore failure and an SDMA page fault under memory pressure, "the process does not crash. Instead every subsequent GPU computation in that process silently returns garbage" — output being pure black or saturated white. The trigger is this allocator configuration:
PYTORCH_ALLOC_CONF=garbage_collection_threshold:0.8,max_split_size_mb:128,expandable_segments:True
The reported workaround is to remove expandable_segments:True and keep everything else identical. The reporter notes this eliminates the problem even though it results in higher peak VRAM usage — which is the evidence that the feature itself, not memory exhaustion, is the cause.
If black frames on an AMD card started after a ROCm upgrade, check your environment for that variable (it is set by a lot of copy-pasted optimisation guides) and check the kernel log for a KFD or SDMA fault. Broader AMD setup is covered in our AMD ROCm local AI setup guide.
Cause 9: What If Every Precision Is Black?
Tell: you have tried fp16, bf16 and fp32 on the same checkpoint and all three come out black, while a sibling build of the same model is fine. At that point you are not debugging your machine any more. Two current reports define the shape of this one, and both are quoted here by title rather than by issue number (see the citation note near the top):
- "Z-Image Base outputs silent black images regardless of precision (FP16/BF16/FP32), while Turbo works perfectly." Precision-independence is the whole diagnosis: cause 1 is a precision overflow, cause 7 is a dispatch difference between two precisions, and both of those predict that some precision works. When none does and the Turbo variant of the same family renders normally, the fault is in the Base weight path, not in your flags. Search the tracker for the live thread before you spend an evening on it. Our Z-Image Base and Z-Image-Edit guide covers what the Base model needs when it does run, and the Z-Image Turbo ComfyUI walkthrough is the variant to fall back to meanwhile.
- "Wan 2.2 official Image-to-Video workflow generates 100% NaN latent tensors on RTX 5070 Ti" — reported against the official models on a clean install, and reproducible. This is the most demoralising version of the failure, because every variable you would normally change is already at its default: official workflow, official weights, nothing custom in the graph. The value of knowing about it is purely that it stops you rebuilding a working install. Search the tracker for the live thread, and check our Wan VRAM requirements by GPU to confirm the build you picked is the one that fits the card before you assume the bug is yours.
What to actually do when you land here. Confirm the sibling build works (Turbo instead of Base; the vendor's repack instead of the plain weights) — that single test converts "my install is broken" into "this checkpoint has a bug", which is a much cheaper thing to live with. Then add your configuration to the existing thread rather than opening a new one: GPU, driver, torch build, ComfyUI version, the exact filename and its size in bytes. Model-side bugs get fixed when the maintainers can see which hardware and which file, and a black frame with no console output tells them nothing on its own.
Which Models Go Black on Which Precision and Backend?
This is the lookup table. Find the row that matches the model, the precision and the machine you are on; the rest of the page is the explanation behind it. Rows marked "quoted by title" are the reports we could not attach a verified issue number to.
| Model / build | Precision | Backend | Symptom | What works |
|---|---|---|---|---|
| SD 1.5 checkpoints | fp16 | Any CUDA | Grey or brown, rarely true black | Load vae-ft-mse-840000; true black here is usually a custom node |
| SDXL / Pony / Illustrious | fp16 VAE | Any CUDA | Pure black only after decode | sdxl-vae-fp16-fix, or --fp32-vae / --bf16-vae |
SDXL with --fast | fp8 ops | Any CUDA | Black since the flag was added (#4572) | Fixed by PR #5928; update, or drop --fast |
| FLUX dev fp8 e4m3fn | fp8 | RTX 4090 (Ada) | Pure black, "invalid value encountered in cast" (#4673, open) | The fp16/bf16 build, or a GGUF quant |
| Kontext Dev via INT4 backend | int4 | RTX 3080 Ti | KSampler itself emits NaN (#9046, open) | No fix posted; move off the INT4 backend |
| Z-Image Turbo | fp16 / bf16 | Any | Renders normally | — |
| Z-Image Base | fp16, bf16 and fp32 | Any | Silent black at every precision (quoted by title) | Use Turbo until the thread closes |
| MiniMax H3 ref2va bf16 (61.7 GB) | bf16 | RTX PRO 6000 Blackwell, v0.32.0 | 100% black frames (#15563, open) | The int8_convrot repack |
| MiniMax H3 pruned bf16 (37.5 GB) | bf16 | Same as above | 100% black frames — pruning changes nothing (#15563) | The int8_convrot repack |
| MiniMax H3 INT8 ConvRot VAE | int8 | — | Black video, NaN/Inf at the VAE (quoted by title) | Decode at higher precision; test halves separately |
| MiniMax H3, any quant | fp8 / int8 / bf16 | RX 7900 XTX | Pure noise, not black (#15314) | Different bug — noise is not NaN |
| Wan 2.2 official I2V | Official weights | RTX 5070 Ti, clean install | 100% NaN latent tensors (quoted by title) | Nothing local to change; track the thread |
| LTX-2.x | bf16 | Apple Silicon / MPS | Random all-NaN attention, black video (quoted by title) | --force-fp32; see the Mac page |
| Boogu T2I Turbo, warm reuse | fp32 VAE already on | RX 7600 XT, ROCm 7.14, dynamic VRAM | First run fine, every reuse black (#15452, open) | Disable dynamic VRAM; unload between runs |
| Krea 2 + WeightHook nodes | bf16 | NVIDIA, async offload on | Black only with hook nodes in the graph (#15531, open) | LoraLoaderModelOnly; disable async offload |
| Any model | Any | RX 9070 XT, ROCm 7.14.0 | Black and white frames after a KFD/SDMA fault (ROCm #6603) | Drop expandable_segments:True |
Read down the Precision column and the 2026 pattern is obvious: the old rule was "more bits are safer", and on MiniMax H3 the 61.7 GB bf16 file is the one that fails while an int8 repack of the same weights renders. Read down the Backend column and the second pattern appears: the same model can be black on one vendor's stack and merely noisy on another's. Neither pattern is something you can flag your way out of — which is why the first move is always to change the file, not the launch arguments.
For the older families, the shorter version still holds. SDXL and its descendants: load sdxl-vae-fp16-fix, then --fp32-vae, and remove --fast if an old tutorial put it in your launch script — see SDXL vs FLUX locally for which family you are actually on. FLUX and FLUX.2: swap the fp8 build for fp16, bf16 or a GGUF, and confirm you loaded the FLUX ae autoencoder rather than an SD VAE.
The Flags, Verified
Read from comfy/cli_args.py on master (ComfyUI v0.33.1). Help text is quoted verbatim from the file:
| Flag | Help text | When to use it |
|---|---|---|
--fp32-vae | "Run the VAE in full precision fp32." | First thing to try for any black output |
--bf16-vae | "Run the VAE in bf16." | Middle ground; bf16 has fp32's exponent range |
--fp16-vae | "Run the VAE in fp16, might cause black images." | Never, when debugging this |
--cpu-vae | "Run the VAE on the CPU." | Diagnostic: isolates decode from sampling |
--force-fp32 | "Force fp32 (If this makes your GPU work better please report it)." | Apple Silicon; last-resort sledgehammer |
--disable-smart-memory | "Force ComfyUI to agressively offload to regular ram instead of keeping models in vram when it can." | Rules out warm-model reuse (cause 5) |
--reserve-vram | "Set the amount of vram in GB you want to reserve for use by your OS/other software..." | When fp32 decode pushes you into OOM |
--async-offload | "Use async weight offloading... Enabled by default on Nvidia." | Disable when testing hook/LoRA black output |
Note the difference between --bf16-vae and --fp16-vae. Both are 16-bit and use the same memory, but bf16 carries the same exponent range as fp32 — it trades mantissa bits for range. That is exactly the tradeoff the SDXL VAE overflow needs, which makes --bf16-vae a genuinely useful middle option when --fp32-vae is too expensive on a small card.
Honest Limitations
- We did not reproduce these on our own hardware. Everything attributed above comes from the linked GitHub issues and model cards, and we say which is which. We are not going to invent a black-image reproduction on an RTX PRO 6000, an RTX 5070 Ti or an M3 Ultra we do not own. Every file size, GPU and version string in the tables came from a reporter's own environment block, not from us.
- Four reports are cited by title, not by number. The Z-Image Base precision-independent case, the MiniMax H3 INT8 ConvRot VAE case, the MPS all-NaN attention case and the Wan 2.2 NaN-latent case are quoted as their tracker titles with a search link, because we could not confirm a stable issue number for them. Treat those four rows as "someone else is hitting this too", not as a citation you can paste.
- Causes 5 through 9 are open bugs and will move. #15452, #15531, #15563 and ROCm #6603 were open when this was written. Click through before you rebuild a workflow around a workaround — the fix may already have shipped.
- Flag names drift.
comfy/cli_args.pyhas changed substantially across the v0.3x releases. The table above is v0.33.1. Runpython main.py --helpand trust your own build over any article, including this one. - We have not published a per-GPU fp8 kernel matrix, because verifying one properly would mean owning every architecture, and we do not. The fp16-vs-fp8 file swap on your own card answers the question for you in ten minutes and is not a guess.
- Custom nodes are a real cause we cannot enumerate. A node that upcasts, downcasts or normalises a latent in the middle of your graph can manufacture NaN on its own. If nothing here fits, retest the same workflow on a clean ComfyUI install with no custom nodes; if that works, bisect your node packs.
FAQ
Why does ComfyUI generate a black image with no error message?
Because NaN is a valid floating-point value, not a crash. The sampler emits it, the VAE decodes it, NumPy clips it to zero when casting to 8-bit, and SaveImage writes a perfectly valid all-black PNG. The only trace is usually a single RuntimeWarning: invalid value encountered in cast line in the console, which is a warning rather than an error and scrolls away. Start with --fp32-vae.
Does --fp32-vae fix every black image?
No, but it is the right first move because it is one flag and it eliminates the most common cause outright. If it does not help, the NaN is being produced during sampling rather than during decode — prove that with the --cpu-vae test, then work through the quantisation, offload and platform causes above.
My image is grey or brown, not black. Same problem?
No. Pure #000000 across every pixel is NaN. A flat grey or muddy brown with faint structure is the wrong VAE for the checkpoint's latent space — SD latents decoded with an SDXL VAE, or FLUX latents decoded with anything that is not the FLUX autoencoder. Load the family's own VAE explicitly with a VAELoader node instead of relying on whatever is baked into the checkpoint.
The first image is fine and every one after it is black. Why?
That is the signature of ComfyUI's memory manager handing back a warm model in a bad state. Issue #15452 documents it with dynamic VRAM: a fresh load succeeds, a reused resident model produces NaN at VAE decode. Restart without the dynamic-VRAM flag, or unload models between generations, and check whether the issue has been closed since.
I tried fp16, bf16 and fp32 and all three are black. What now?
Then it is not a precision problem, and no flag on this page will help. Precision-independence is the signature of a model-side bug — the open Z-Image Base report is exactly this, black at FP16, BF16 and FP32 while Z-Image Turbo renders normally on the same install. Test the sibling build (Turbo instead of Base, or the vendor's quantised repack instead of the plain weights). If the sibling works, you have proven your install is fine and the checkpoint is not, which is the cheapest possible outcome.
My images are fine but my videos come out black. Is that the same bug?
Usually not the same instance, though it is the same failure mode. The 2026 video checkpoints are where most of the current reports sit: MiniMax H3 bf16 giving 100% black frames while its int8 repack works (#15563), an INT8 ConvRot VAE emitting NaN/Inf, LTX-2.x hitting random all-NaN attention on Apple Silicon, and a Wan 2.2 Image-to-Video workflow producing 100% NaN latents on an RTX 5070 Ti. Find your model in the lookup table above rather than assuming the image-model fix carries over — on video models the quantised build is often the one that works, which is the opposite of the SDXL-era advice.
Should I use --fp16-vae to make the VAE faster?
Not while you are debugging a black image. ComfyUI's own help text for that flag is "Run the VAE in fp16, might cause black images." If you want the memory savings, use --bf16-vae instead — same 16 bits, but bf16 keeps fp32's exponent range, which is the specific thing the SDXL VAE overflows.
Is a black image ever a prompt or sampler problem?
Almost never a pure black one. Extreme CFG values, a broken negative embedding or a badly mismatched sampler/scheduler pair produce ugly, burnt or noisy images — not a uniform zero frame. The 1-step, CFG-1.0 test at the top of this page separates the two in a minute: if a 1-step generation is coloured and a 20-step generation is black, something numerical is going wrong during sampling, and no amount of prompt editing will touch it.
Sources
- ComfyUI — comfy/cli_args.py (flag names and help text, master, v0.33.1 released 13 Aug 2026)
- madebyollin/sdxl-vae-fp16-fix (fp16 NaN explanation and the rescaled VAE)
- Comfy-Org/ComfyUI issues #4572, #4673, #9046, #10681, #15314, #15315, #15452, #15531, #15563, #15617, and PR #5928
- Cited by tracker title, not by number (search links): Z-Image Base black at every precision, MiniMax H3 INT8 ConvRot VAE NaN/Inf, MPS all-NaN attention on bf16 LTX-2.x, Wan 2.2 I2V 100% NaN latents
- ROCm/ROCm issue #6603 (silent zero readback on gfx1201)
Generating images locally? Take it further.
From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.
Go from one-off images to a real workflow
The Local Image Generation course covers ComfyUI, SDXL and FLUX properly — plus 24 more courses on running AI on your own hardware.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARRun FLUX.1 Locally in 2026: VRAM Needs + 5-Minute Setup
- AI-Toolkit LoRA Training: FLUX.2, Z-Image & Qwen-Image
- Best GPU for Local AI Image Generation (2026): Ranked
- Best Local AI Image Models 2026: FLUX vs SDXL vs Qwen
- blog/flux-vram-requirements-by-gpu
- Chroma Local Guide: The Apache-2.0 Uncensored FLUX Model
- ComfyUI FLUX Workflow (2026): JSON Nodes Explained
- ComfyUI IMPORT FAILED: Find the Real Error Fast
- ComfyUI LoRA Not Working: Key Not Loaded Fixes
- ComfyUI Manager Install Failed: Registry and Path Fixes
Comments (0)
No comments yet. Be the first to share your thoughts!