★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Image Generation

ComfyUI Black Image Fix: NaN, VAE and fp8 by Model

September 20, 2026
18 min read
Local AI Master Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Generating images locally? Take it further. From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.

Start free
Or own it for life — Lifetime $149, pay once

A pure black output with no error message means your latent went to NaN before the VAE decoded it, and on most machines the single fastest fix is to restart ComfyUI with --fp32-vae. That is not folklore: ComfyUI's own argument parser describes the opposite flag, --fp16-vae, as "Run the VAE in fp16, might cause black images." If --fp32-vae does not fix it, the cause is specific to the model, the precision and the backend you are running: an fp16-overflowing SDXL VAE, an fp8 or int-quantised checkpoint without the matching kernels, a bf16 weight path on a brand-new 2026 checkpoint, a warm model handed back by the memory manager, or a checkpoint that comes out black at every precision you try. Grey or brown output is a different bug entirely — that is the wrong VAE, not a NaN.

The reason this failure is so hard to search is that nothing fails. The queue completes, the progress bar fills, the node graph goes green, and the SaveImage node writes a valid PNG that happens to be 100% #000000. There is no traceback to paste. This page gives you the thread to pull: a two-node test that proves whether it is NaN, then the nine causes ordered by how often each one turns out to be the real one, then a model / precision / backend lookup table you can scan straight to your own row — because the fix for SDXL actively makes the FLUX case worse, and the 2026 model releases inverted the rule that "less quantisation is safer".

Flag names below were read from comfy/cli_args.py on master at the time of writing (ComfyUI v0.33.1, released 13 August 2026 — that file has changed a lot across the v0.3x series, so check yours with python main.py --help). Bug reports are linked to their issue numbers so you can check whether yours has been fixed since.

A note on how the newest reports are cited. Four of the 2026 cases below are quoted by their tracker title rather than by a number, because we could not confirm a stable issue number for them. Where that is the case we say so and link a tracker search for the title instead of a specific issue — a search that finds the live thread is more useful to you than a number we are not certain of.

What Error Message Should I Be Looking For?

The black image itself produces no error. But in a huge share of these reports there is one line in the console, and it is the giveaway:

RuntimeWarning: invalid value encountered in cast
  img = Image.fromarray(np.clip(i, 0, 255).astype(np.uint8))

That warning appears in issue #4673 (FLUX dev fp8 outputting pure black on a 4090) and again in issue #10681 (black images on an M3 Ultra Mac Studio). It means NumPy was handed NaN and cast it to an integer. NaN clips to zero. Zero is black.

If you see that line, stop reading about prompts, samplers, CFG and negative embeddings. Your pipeline produced not-a-number and the rest of this page applies. If you do not see it, scroll up in the console anyway — it is a warning, not an error, so it prints once and gets buried by the next queue item.

Coming from Automatic1111 or Forge instead? The equivalent knob there is --no-half-vae, and the same underlying overflow is what it works around.

Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

How Do I Confirm It Is a NaN and Not a Bad Prompt?

Two nodes settle it in under a minute, and you should do this before changing anything.

  1. Bypass the VAE. Wire your KSampler's LATENT output into a PreviewImage through a fresh VAEDecode, but first drop the sampler to 1 step and set CFG to 1.0. A 1-step, CFG-1 generation should produce a blurry, ugly, but coloured image. If 1 step is coloured and 20 steps is black, the NaN is being produced during sampling — cause 3, 5, 6 or 7 below. If even 1 step is black, the decode is the problem — cause 1 or 2.
  2. Move the decode to CPU. Restart with --cpu-vae (the help text is simply "Run the VAE on the CPU."). CPU tensors are fp32 by default. If the image comes back correctly on the CPU VAE, you have proven a GPU-side precision overflow and --fp32-vae is your permanent fix. If it is still black on the CPU, the latent arriving at the VAE was already NaN and the VAE is innocent.

That second test is the most useful diagnostic on this page, because it cleanly splits "the decode broke" from "the sample broke" — and every fix below sits on one side or the other of that line.

What Causes a Black Image in ComfyUI? The Nine, Ranked

#CauseTellFix
1SDXL-family VAE overflows in fp16Black only after decode; CPU VAE works--fp32-vae, or load sdxl-vae-fp16-fix
2Wrong or missing VAE for the checkpointOutput is grey/brown/washed, not pure blackLoad the VAE the model was trained with
3fp8 / int-quantised weights without matching kernelsBlack on one GPU, fine on another, same fileUse the fp16/bf16 build, or a GGUF quant
4Apple Silicon MPS precisionBlack on a Mac, same graph fine on CUDA--force-fp32, or --cpu-vae
5Warm model reuse under dynamic VRAMFirst run good, every later run blackRestart without dynamic VRAM; unload between runs
6Hook/LoRA weights lost after offloadBlack only when a hook node is in the graphUse LoraLoaderModelOnly; disable async offload
7bf16 path on a brand-new checkpointBlack on bf16 file, fine on the quantised repackUse the vendor's quantised build for now
8AMD ROCm allocator returning zerosBlack and white frames, after a KFD fault in dmesgDrop expandable_segments:True
9The checkpoint is black at every precisionfp16, bf16 and fp32 all black; a sibling build worksModel-side bug: use the sibling build, report yours

Work down that list in order. Causes 1 and 2 account for the overwhelming majority of what people are actually hitting; 5 through 9 are current-release bugs that will churn — and cause 9 is the one that did not exist before the late-2026 model wave, when several vendors shipped a fast variant that works and a full variant that does not.

Cause 1: The SDXL fp16 VAE Overflow (Start Here)

This is the classic, and it is a property of the original SDXL VAE weights, not of your install. The sdxl-vae-fp16-fix model card states the problem in one sentence: "SDXL-VAE generates NaNs in fp16 because the internal activation values are too big." fp16's maximum representable value is 65,504; the decoder's internal activations exceed it, the tensor goes to infinity, infinity minus infinity is NaN, and everything downstream is NaN.

Two fixes, and they are not equivalent:

Fix A — run the VAE in fp32 (universal, costs VRAM):

python main.py --fp32-vae

This works for every model, not just SDXL, because it removes the overflow condition entirely. The cost is real: the decode step holds fp32 activations, which on a tight 8GB card can push you from a black image into an out-of-memory error instead. If that happens, see our guide to ComfyUI's memory manager and dynamic VRAM for the offloading knobs — trading one failure for another is progress, but only just.

Fix B — use the rescaled VAE (cheaper, SDXL only): download sdxl.vae.safetensors from madebyollin/sdxl-vae-fp16-fix into ComfyUI/models/vae/, add a VAELoader node, and wire it into your VAEDecode instead of the checkpoint's built-in VAE. The author fine-tuned the VAE to produce the same final output with smaller internal activations by "scaling down weights and biases within the network," so it simply never reaches the overflow. The model card is honest about the tradeoff: "There are slight discrepancies between the output of SDXL-VAE-FP16-Fix and SDXL-VAE, but the decoded images should be close enough for most purposes."

Use Fix B if you generate SDXL all day and want your VRAM back. Use Fix A if you switch between model families, because Fix B does nothing for FLUX, Qwen-Image or a video model.

One warning that applies to both: do not "fix" this by adding --fp16-vae. People find that flag while searching and try it because it looks like a precision setting. It is the setting that causes the bug, which is why its own help text says so.

Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

Cause 2: Grey, Brown or Washed-Out Is a VAE Mismatch, Not a NaN

If your output is a flat grey or muddy brown rather than true black, stop reading about NaN — you decoded the latent with the wrong VAE. This is the second-most-common report and it gets misfiled under "black image" constantly, because at a glance a dark grey 512×512 square looks black in a file browser thumbnail.

Every model family uses its own latent space. SD 1.5 latents decoded with an SDXL VAE, or FLUX latents decoded with an SD VAE, produce structured garbage — often a uniform mid-tone with faint colour blocking — because the decoder is interpreting the numbers correctly according to a different mapping. The check: open the PNG in an editor and read the pixel values. Exactly 0,0,0 across the whole frame is NaN. Anything else — 40,38,35 or a smeary brown — is a VAE mismatch.

The fix is to load the matching VAE explicitly rather than relying on whatever is baked into the checkpoint:

  • SD 1.5: the checkpoint's built-in VAE is usually fine; vae-ft-mse-840000-ema-pruned is the common upgrade.
  • SDXL and its descendants: sdxl_vae.safetensors, or the fp16-fix version above.
  • FLUX: the FLUX autoencoder (ae.safetensors) — nothing else will decode a FLUX latent. Our FLUX VRAM requirements guide lists the support files alongside each diffusion model, and the FLUX.2 setup guide covers the current generation's file layout.
  • Qwen-Image / Z-Image and the video models: each ships its own VAE file. Do not reuse one across families. The Z-Image Turbo ComfyUI guide walks through the correct loader wiring for that family.

A related trap: some checkpoints are distributed without a baked VAE. ComfyUI will happily run the graph and hand you the mismatch rather than erroring, so a missing VAE and a wrong VAE look identical from the outside.

Cause 3: fp8 and Int-Quantised Checkpoints Without the Kernels

The symptom that identifies this one: the exact same file produces a good image on one GPU and a black image on another. Quantised weights are only half the story — the other half is a kernel that can multiply them, and that kernel is compiled for specific architectures.

The history in ComfyUI's own tracker is instructive:

  • #4572, "SDXL generate black images with new --fast arg" (August 2024) was fixed by PR #5928, titled "Fix SDXL generating black images when using --fast (fp8_ops)". The --fast path swapped in fp8 matmuls that SDXL's numerics could not survive.
  • #4673, "The flux dev model outputs pure black on 4090", reports the fp8 e4m3fn FLUX dev build going black on an RTX 4090 while FLUX Schnell and other models are fine on the same machine, with the "invalid value encountered in cast" warning in the log. It is still open.
  • #9046, "Black images in XL and Nunchuka(Kontext Dev) caused by Ksampler outputting NaNs. Nothing works.", ties the same failure to an INT4 inference backend on an RTX 3080 Ti. It is still open.

What to do about it, in order:

  1. Swap the file, not the flags. If the fp8 build is black, download the fp16 or bf16 build of the same model and confirm it works. That single test tells you whether you are looking at a quantisation problem or something else, and it costs you a download rather than an afternoon.
  2. If fp16 works and fp8 does not, use a GGUF quant instead. GGUF quants dequantise to a dtype your card definitely supports rather than relying on native fp8 tensor cores. You give up some speed and get correctness.
  3. Do not add --fast. If you have it in your launch script from an old tutorial, remove it and retest — it is precisely the sort of aggressive-op flag that produced #4572.
  4. Check the arch, honestly. Native fp8 matmul is a Blackwell/Ada-class capability; on older cards ComfyUI stores fp8 weights and upcasts them, which is a different code path with different numerics. We do not own one of every GPU generation and we are not going to publish a card-by-card kernel matrix we cannot verify — the reliable test is the fp16-vs-fp8 swap above, run on your GPU. If the quantised build is failing because it will not fit rather than because it will not compute, our low-VRAM FLUX guide covers the offload options that keep you on a dtype your card actually supports.

The explicit dtype flags exist if you want to force the issue during debugging: --fp8_e4m3fn-unet, --fp8_e5m2-unet, --fp16-unet, --bf16-unet, --fp32-unet, plus the text-encoder equivalents --fp8_e4m3fn-text-enc, --fp16-text-enc and --fp32-text-enc. Forcing --fp16-unet on a model that expects bf16 is itself a way to manufacture a black image, so change one thing at a time.

Cause 4: Why Do I Only Get Black Images on a Mac?

On a Mac, try --force-fp32 before anything else, then --fp32-vae, then --cpu-vae. ComfyUI's own log line when it engages the first of those reads "Forcing FP32, if this improves things please report it," which tells you how the maintainers regard MPS numerics. Each step costs memory and speed, and on unified memory that cost is felt immediately.

Two things are worth knowing before you start swapping flags. First, issue #10681 — black images on a Mac Studio (M3 Ultra, 256GB) with a WanVAE / Qwen-Image graph while the same workflow is correct on an M3 Max MacBook Pro — reproduces in a separate app too, so it is an MPS-level problem rather than a ComfyUI bug. Second, the 2026 video models added a distinctly different Mac symptom: an open tracker report describes random all-NaN attention on MPS producing black videos with bf16 LTX-2.x models, and proposes two fixes in-thread (search the tracker — this is one of the reports we are quoting by title, not by number). The report's own word is "random", so intermittency is the signature — a graph that renders correctly on one run and black on the next, with nothing changed, is this and not a precision setting you got wrong. If you are on that family, our LTX-2 local setup guide has the working file layout to compare against.

Everything else Mac-specific lives on its own page. Rather than duplicate it here, see ComfyUI on Mac: MPS errors and what fixes them for the backend errors, the fallback behaviour and the attention-implementation flags. (Historic note for searchers: #15315, the official MiniMax H3 text-to-video workflow producing black video and NaN audio on an M4 Max, has been closed.)

Cause 5: Warm Models Under Dynamic VRAM

Tell: the first generation after a fresh model load is perfect, and every subsequent generation from the same queue is black. That pattern is unmistakable once you know to look for it, and it is a live regression rather than a configuration mistake.

Issue #15452 — "Dynamic VRAM: reused (warm) model produces NaN/black output on VAE decode, fresh load does not", opened 9 August 2026 and open at the time of writing — reports that when prompts are queued back to back against the Boogu T2I Turbo official template, "every subsequent generation that reuses the already-resident Boogu model renders fully black", with the same "invalid value encountered in cast" warning. The reporter's environment was ComfyUI v0.31.0, an RX 7600 XT on ROCm 7.14, with dynamic VRAM enabled and the VAE already running in fp32 — note that fp32 VAE was already on, which is what makes this a distinct cause rather than cause 1 in disguise. Their evidence is a missing log line: runs that printed "Requested to load Boogu" succeeded, and runs that skipped straight to the dynamic-VRAM prepare step failed.

Workarounds, all unsatisfying: restart without the dynamic-VRAM flag; or unload models between generations instead of queueing prompts back to back, which throws away most of what dynamic VRAM is for. Check whether the issue has been closed before you rearrange your workflow around it.

If you arrived here from an out-of-memory error after an update rather than a black image, that is the same subsystem behaving differently and a different diagnosis — the memory-manager section of our complete ComfyUI guide is the better starting point.

Cause 6: Hooks and LoRAs That Vanish After Offload

Tell: the graph is black only when a hook node is present, and identical without it. Issue #15531, "BF16 WeightHooks can produce black output after model offload" (opened 12 August 2026, open), records a mapped WeightHook producing an all-black image on a BF16 Krea 2 checkpoint even though the prompt completes successfully.

The decisive detail for diagnosis: with DynamicVRAM enabled, the same LoRA works through LoraLoaderModelOnly and, through CreateHookLoraModelOnly → SetClipHooks, completes but yields what the reporter calls "black/non-finite output". Their diagnosis is that hook writeback after model preparation and offload has to preserve a valid storage and transfer lifecycle, and the candidate fix is to keep plain-weight hook transfers synchronous rather than pinned-async.

So: if your graph uses hook nodes, rebuild that section with the plain LoRA loader and retest. If that fixes it, you have confirmed the cause and can decide whether you need hooks badly enough to also disable async offload (--async-offload controls the number of offload streams and is on by default on NVIDIA).

Cause 7: bf16 Paths on the Newest Checkpoints

Tell: the vendor's quantised repack works and the plain bf16 file of the same weights does not. This is the 2026-specific version of the problem, and it shows up within days of each new model release.

Issue #15563 (opened 13 August 2026, open) is the cleanest example on record. The reporter tested three files of the same MiniMax H3 weights on ComfyUI v0.32.0, Windows 11, an RTX PRO 6000 Blackwell 96GB, torch 2.12.0+cu130:

CheckpointSizeResult
minimax_h3_ref2va_bf16.safetensors61.7 GB100% black, every frame
minimax_h3_ref2va_pruned_bf16.safetensors37.5 GB100% black, every frame
minimax_h3_ref2va_int8_convrot.safetensors—Correct output

Their conclusion, after ruling out file corruption, attention backends, VAE precision and memory pressure: "the only path that differs between the working and failing runs is runtime op dispatch: int8_convrot layers execute through the comfy_kitchen quantized kernels, while plain bf16 layers execute through standard ops."

Note what this inverts. Everywhere else on this page, quantisation is the suspect and full precision is the safe fallback. Here it is the reverse — the quantised build is the one that works. The practical rule for any checkpoint released in the last few weeks: if one precision variant is black, try the others before you conclude anything about your hardware. Vendors ship several builds precisely because the fast paths are new code.

But do not stop at the transformer. The int8 repack that fixes the sampler has its own reported failure at the other end of the graph: a separate open report, titled in the tracker as the MiniMax H3 INT8 ConvRot VAE producing black video with NaN/Inf output, puts the non-finite values at the VAE stage rather than during sampling (quoted by title — search the tracker for the live thread). The practical consequence is that on this family the two halves of the pipeline want different things, so test them separately: run the --cpu-vae check from the top of this page, and if the quantised transformer plus a higher-precision VAE decode gives you a picture, keep that pairing. It is the same split the diagnostic at the top exists to expose, just on a model where both sides have live bugs at once.

Two adjacent reports from the same wave, worth knowing about so you do not misfile your own symptom: #15314 has MiniMax H3 producing pure noise (not black) on an RX 7900 XTX across every quantisation and backend combination, and #15617 has an INT8 ConvRot build hard-rebooting a Windows machine on an RTX 3080. Noise is not NaN, and a reboot is not a numerical failure.

Cause 8: AMD, ROCm and Silent Zeros

On an AMD card, check dmesg before you touch ComfyUI. There is a ROCm-level failure mode that produces exactly the symptom this page is about, and no amount of VAE precision will fix it.

ROCm issue #6603 (RX 9070 XT, gfx1201, ROCm 7.14.0, a regression from 7.2.4) reports that after a KFD userptr restore failure and an SDMA page fault under memory pressure, "the process does not crash. Instead every subsequent GPU computation in that process silently returns garbage" — output being pure black or saturated white. The trigger is this allocator configuration:

PYTORCH_ALLOC_CONF=garbage_collection_threshold:0.8,max_split_size_mb:128,expandable_segments:True

The reported workaround is to remove expandable_segments:True and keep everything else identical. The reporter notes this eliminates the problem even though it results in higher peak VRAM usage — which is the evidence that the feature itself, not memory exhaustion, is the cause.

If black frames on an AMD card started after a ROCm upgrade, check your environment for that variable (it is set by a lot of copy-pasted optimisation guides) and check the kernel log for a KFD or SDMA fault. Broader AMD setup is covered in our AMD ROCm local AI setup guide.

Cause 9: What If Every Precision Is Black?

Tell: you have tried fp16, bf16 and fp32 on the same checkpoint and all three come out black, while a sibling build of the same model is fine. At that point you are not debugging your machine any more. Two current reports define the shape of this one, and both are quoted here by title rather than by issue number (see the citation note near the top):

  • "Z-Image Base outputs silent black images regardless of precision (FP16/BF16/FP32), while Turbo works perfectly." Precision-independence is the whole diagnosis: cause 1 is a precision overflow, cause 7 is a dispatch difference between two precisions, and both of those predict that some precision works. When none does and the Turbo variant of the same family renders normally, the fault is in the Base weight path, not in your flags. Search the tracker for the live thread before you spend an evening on it. Our Z-Image Base and Z-Image-Edit guide covers what the Base model needs when it does run, and the Z-Image Turbo ComfyUI walkthrough is the variant to fall back to meanwhile.
  • "Wan 2.2 official Image-to-Video workflow generates 100% NaN latent tensors on RTX 5070 Ti" — reported against the official models on a clean install, and reproducible. This is the most demoralising version of the failure, because every variable you would normally change is already at its default: official workflow, official weights, nothing custom in the graph. The value of knowing about it is purely that it stops you rebuilding a working install. Search the tracker for the live thread, and check our Wan VRAM requirements by GPU to confirm the build you picked is the one that fits the card before you assume the bug is yours.

What to actually do when you land here. Confirm the sibling build works (Turbo instead of Base; the vendor's repack instead of the plain weights) — that single test converts "my install is broken" into "this checkpoint has a bug", which is a much cheaper thing to live with. Then add your configuration to the existing thread rather than opening a new one: GPU, driver, torch build, ComfyUI version, the exact filename and its size in bytes. Model-side bugs get fixed when the maintainers can see which hardware and which file, and a black frame with no console output tells them nothing on its own.

Which Models Go Black on Which Precision and Backend?

This is the lookup table. Find the row that matches the model, the precision and the machine you are on; the rest of the page is the explanation behind it. Rows marked "quoted by title" are the reports we could not attach a verified issue number to.

Model / buildPrecisionBackendSymptomWhat works
SD 1.5 checkpointsfp16Any CUDAGrey or brown, rarely true blackLoad vae-ft-mse-840000; true black here is usually a custom node
SDXL / Pony / Illustriousfp16 VAEAny CUDAPure black only after decodesdxl-vae-fp16-fix, or --fp32-vae / --bf16-vae
SDXL with --fastfp8 opsAny CUDABlack since the flag was added (#4572)Fixed by PR #5928; update, or drop --fast
FLUX dev fp8 e4m3fnfp8RTX 4090 (Ada)Pure black, "invalid value encountered in cast" (#4673, open)The fp16/bf16 build, or a GGUF quant
Kontext Dev via INT4 backendint4RTX 3080 TiKSampler itself emits NaN (#9046, open)No fix posted; move off the INT4 backend
Z-Image Turbofp16 / bf16AnyRenders normally—
Z-Image Basefp16, bf16 and fp32AnySilent black at every precision (quoted by title)Use Turbo until the thread closes
MiniMax H3 ref2va bf16 (61.7 GB)bf16RTX PRO 6000 Blackwell, v0.32.0100% black frames (#15563, open)The int8_convrot repack
MiniMax H3 pruned bf16 (37.5 GB)bf16Same as above100% black frames — pruning changes nothing (#15563)The int8_convrot repack
MiniMax H3 INT8 ConvRot VAEint8—Black video, NaN/Inf at the VAE (quoted by title)Decode at higher precision; test halves separately
MiniMax H3, any quantfp8 / int8 / bf16RX 7900 XTXPure noise, not black (#15314)Different bug — noise is not NaN
Wan 2.2 official I2VOfficial weightsRTX 5070 Ti, clean install100% NaN latent tensors (quoted by title)Nothing local to change; track the thread
LTX-2.xbf16Apple Silicon / MPSRandom all-NaN attention, black video (quoted by title)--force-fp32; see the Mac page
Boogu T2I Turbo, warm reusefp32 VAE already onRX 7600 XT, ROCm 7.14, dynamic VRAMFirst run fine, every reuse black (#15452, open)Disable dynamic VRAM; unload between runs
Krea 2 + WeightHook nodesbf16NVIDIA, async offload onBlack only with hook nodes in the graph (#15531, open)LoraLoaderModelOnly; disable async offload
Any modelAnyRX 9070 XT, ROCm 7.14.0Black and white frames after a KFD/SDMA fault (ROCm #6603)Drop expandable_segments:True

Read down the Precision column and the 2026 pattern is obvious: the old rule was "more bits are safer", and on MiniMax H3 the 61.7 GB bf16 file is the one that fails while an int8 repack of the same weights renders. Read down the Backend column and the second pattern appears: the same model can be black on one vendor's stack and merely noisy on another's. Neither pattern is something you can flag your way out of — which is why the first move is always to change the file, not the launch arguments.

For the older families, the shorter version still holds. SDXL and its descendants: load sdxl-vae-fp16-fix, then --fp32-vae, and remove --fast if an old tutorial put it in your launch script — see SDXL vs FLUX locally for which family you are actually on. FLUX and FLUX.2: swap the fp8 build for fp16, bf16 or a GGUF, and confirm you loaded the FLUX ae autoencoder rather than an SD VAE.

The Flags, Verified

Read from comfy/cli_args.py on master (ComfyUI v0.33.1). Help text is quoted verbatim from the file:

FlagHelp textWhen to use it
--fp32-vae"Run the VAE in full precision fp32."First thing to try for any black output
--bf16-vae"Run the VAE in bf16."Middle ground; bf16 has fp32's exponent range
--fp16-vae"Run the VAE in fp16, might cause black images."Never, when debugging this
--cpu-vae"Run the VAE on the CPU."Diagnostic: isolates decode from sampling
--force-fp32"Force fp32 (If this makes your GPU work better please report it)."Apple Silicon; last-resort sledgehammer
--disable-smart-memory"Force ComfyUI to agressively offload to regular ram instead of keeping models in vram when it can."Rules out warm-model reuse (cause 5)
--reserve-vram"Set the amount of vram in GB you want to reserve for use by your OS/other software..."When fp32 decode pushes you into OOM
--async-offload"Use async weight offloading... Enabled by default on Nvidia."Disable when testing hook/LoRA black output

Note the difference between --bf16-vae and --fp16-vae. Both are 16-bit and use the same memory, but bf16 carries the same exponent range as fp32 — it trades mantissa bits for range. That is exactly the tradeoff the SDXL VAE overflow needs, which makes --bf16-vae a genuinely useful middle option when --fp32-vae is too expensive on a small card.

Honest Limitations

  • We did not reproduce these on our own hardware. Everything attributed above comes from the linked GitHub issues and model cards, and we say which is which. We are not going to invent a black-image reproduction on an RTX PRO 6000, an RTX 5070 Ti or an M3 Ultra we do not own. Every file size, GPU and version string in the tables came from a reporter's own environment block, not from us.
  • Four reports are cited by title, not by number. The Z-Image Base precision-independent case, the MiniMax H3 INT8 ConvRot VAE case, the MPS all-NaN attention case and the Wan 2.2 NaN-latent case are quoted as their tracker titles with a search link, because we could not confirm a stable issue number for them. Treat those four rows as "someone else is hitting this too", not as a citation you can paste.
  • Causes 5 through 9 are open bugs and will move. #15452, #15531, #15563 and ROCm #6603 were open when this was written. Click through before you rebuild a workflow around a workaround — the fix may already have shipped.
  • Flag names drift. comfy/cli_args.py has changed substantially across the v0.3x releases. The table above is v0.33.1. Run python main.py --help and trust your own build over any article, including this one.
  • We have not published a per-GPU fp8 kernel matrix, because verifying one properly would mean owning every architecture, and we do not. The fp16-vs-fp8 file swap on your own card answers the question for you in ten minutes and is not a guess.
  • Custom nodes are a real cause we cannot enumerate. A node that upcasts, downcasts or normalises a latent in the middle of your graph can manufacture NaN on its own. If nothing here fits, retest the same workflow on a clean ComfyUI install with no custom nodes; if that works, bisect your node packs.

FAQ

Why does ComfyUI generate a black image with no error message?

Because NaN is a valid floating-point value, not a crash. The sampler emits it, the VAE decodes it, NumPy clips it to zero when casting to 8-bit, and SaveImage writes a perfectly valid all-black PNG. The only trace is usually a single RuntimeWarning: invalid value encountered in cast line in the console, which is a warning rather than an error and scrolls away. Start with --fp32-vae.

Does --fp32-vae fix every black image?

No, but it is the right first move because it is one flag and it eliminates the most common cause outright. If it does not help, the NaN is being produced during sampling rather than during decode — prove that with the --cpu-vae test, then work through the quantisation, offload and platform causes above.

My image is grey or brown, not black. Same problem?

No. Pure #000000 across every pixel is NaN. A flat grey or muddy brown with faint structure is the wrong VAE for the checkpoint's latent space — SD latents decoded with an SDXL VAE, or FLUX latents decoded with anything that is not the FLUX autoencoder. Load the family's own VAE explicitly with a VAELoader node instead of relying on whatever is baked into the checkpoint.

The first image is fine and every one after it is black. Why?

That is the signature of ComfyUI's memory manager handing back a warm model in a bad state. Issue #15452 documents it with dynamic VRAM: a fresh load succeeds, a reused resident model produces NaN at VAE decode. Restart without the dynamic-VRAM flag, or unload models between generations, and check whether the issue has been closed since.

I tried fp16, bf16 and fp32 and all three are black. What now?

Then it is not a precision problem, and no flag on this page will help. Precision-independence is the signature of a model-side bug — the open Z-Image Base report is exactly this, black at FP16, BF16 and FP32 while Z-Image Turbo renders normally on the same install. Test the sibling build (Turbo instead of Base, or the vendor's quantised repack instead of the plain weights). If the sibling works, you have proven your install is fine and the checkpoint is not, which is the cheapest possible outcome.

My images are fine but my videos come out black. Is that the same bug?

Usually not the same instance, though it is the same failure mode. The 2026 video checkpoints are where most of the current reports sit: MiniMax H3 bf16 giving 100% black frames while its int8 repack works (#15563), an INT8 ConvRot VAE emitting NaN/Inf, LTX-2.x hitting random all-NaN attention on Apple Silicon, and a Wan 2.2 Image-to-Video workflow producing 100% NaN latents on an RTX 5070 Ti. Find your model in the lookup table above rather than assuming the image-model fix carries over — on video models the quantised build is often the one that works, which is the opposite of the SDXL-era advice.

Should I use --fp16-vae to make the VAE faster?

Not while you are debugging a black image. ComfyUI's own help text for that flag is "Run the VAE in fp16, might cause black images." If you want the memory savings, use --bf16-vae instead — same 16 bits, but bf16 keeps fp32's exponent range, which is the specific thing the SDXL VAE overflows.

Is a black image ever a prompt or sampler problem?

Almost never a pure black one. Extreme CFG values, a broken negative embedding or a badly mismatched sampler/scheduler pair produce ugly, burnt or noisy images — not a uniform zero frame. The 1-step, CFG-1.0 test at the top of this page separates the two in a minute: if a 1-step generation is coloured and a 20-step generation is black, something numerical is going wrong during sampling, and no amount of prompt editing will touch it.

Sources

🎯
AI Learning Path

Generating images locally? Take it further.

From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Once your hardware is sorted

Go from one-off images to a real workflow

The Local Image Generation course covers ComfyUI, SDXL and FLUX properly — plus 24 more courses on running AI on your own hardware.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsComfyUITroubleshootingVAENaNfp8bf16SDXLFLUXZ-ImageWanROCm

Local AI Master Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on Local Image Generation
See the full Run FLUX.1 Locally guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

📅 Published: September 20, 2026🔄 Last Updated: September 20, 2026✓ Manually Reviewed

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once

Was this helpful?

LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Generating images locally? Take it further.

From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators