★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Image Generation

ComfyUI Out of Memory: HostBuffer & DynamicVRAM Fix

September 13, 2026
13 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Generating images locally? Take it further. From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.

Start free
Or own it for life — Lifetime $149, pay once

Short answer: add --disable-dynamic-vram to your launch command and restart. If ComfyUI started OOMing, hanging at "Model Initializing...", or running several times slower on a graph that worked last month, the cause is almost certainly DynamicVRAM — the memory manager ComfyUI switched on by default in February 2026. And --lowvram will not help you, because ComfyUI's own help text for that flag now reads, verbatim: "Doesn't do anything if dynamic vram is enabled."

If your console shows HostBuffer.read_file_slice failed immediately before the CUDA out-of-memory traceback, you have hit a specific, dated regression — not a "your GPU is too small" problem. That is issue #15255, opened 2026-08-03, and with 60 comments it is the most-commented open issue on the ComfyUI tracker. It has its own section below, including the two extra flags multi-GPU reporters needed.

That verbatim --lowvram sentence is why so many people have added the flag, restarted, and watched nothing change.

Everything below was read directly out of comfy/cli_args.py on ComfyUI master on 18 August 2026, pinned against v0.33.1 (released 2026-08-13); the issue tracker was re-read on 23 August 2026. This subsystem changes weekly, so re-verify the flag names against your own install before you trust any article, including this one. There is a one-line command for that in the Measure It Yourself section.


The 30-Second Fix

Append the flag to however you normally launch ComfyUI, then restart the server. Nothing else changes.

# Linux / macOS / manual install
python main.py --disable-dynamic-vram

# Windows portable — edit run_nvidia_gpu.bat and append the flag
# to the existing python line, e.g.
.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --disable-dynamic-vram

Its help string is "Disable dynamic VRAM and use estimate based model loading" — in other words, the pre-2026 behaviour that every guide written before March 2026 silently assumes you are running.

Run your problem workflow again. If the OOM, hang or slowdown is gone, you have your answer and you can stop reading. If it is unchanged, DynamicVRAM was not your problem — jump to Symptom → Cause → Fix, because several of the failures people blame on it are something else wearing its clothes.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

Which Error String Do You Have?

Find the line your console actually printed, not the one you assume it printed. DynamicVRAM produces at least nine distinct failure signatures, and they do not all have the same fix — one of them is cleared by a package downgrade rather than a flag, and two have no fix at all yet. Every row below is an open report on Comfy-Org/ComfyUI, re-read 2026-08-23; issue titles and error strings are quoted, not paraphrased.

The string in your consoleIssueOpenedAlso in playWhat the thread actually does
HostBuffer.read_file_slice failed, then torch.AcceleratorError: CUDA error: out of memory#15255 — 60 comments, the most-commented open issue2026-08-03comfy-aimdo streaming; reported on 0.30.1; reproduces on multi-GPU--disable-dynamic-vram; multi-GPU reporters also used --cuda-device 0 or --disable-pinned-memory
RuntimeError: HostBuffer.truncate failed plus CUDA error: an illegal memory access was encountered#155912026-08-13comfy-kitchen 0.2.31; MiniMaxH3; fails inside partially_unload_ramDowngrade the package: pip install comfy-kitchen==0.2.30
VRAM Allocation failed (non OOM) and Fault failed: 2 on the second generation, never the first#152692026-08-03AIMDO 0.4.11; QwenImageOpen, no maintainer fix — --disable-dynamic-vram is the only lever on your side of the machine
aimdo:VRAM Allocation failed (non OOM) in 17.1 (title verbatim)#129432026-03-14Same fault class, opened months earlier — this one is not newOpen since March; treat a "non OOM" allocation failure as a manager bug, not a capacity problem
618x hostbuf_grow ERROR with RAM climbing across a single long prompt#155752026-08-13AIMDO host buffer; LoRA-loaded workflowCap the cache with --cache-ram 4, or disable the manager outright
--reserve-vram accepted at launch and then silently ignored#156662026-08-16Flag was reworked for DynamicVRAM by PR #14480Use --vram-headroom N instead — it is the flag written for this code path
OOM on Linux from one large single allocation that should fit#14340 — "VRAM OOM on Linux with large singular allocation but should be within limits"2026-06-08Linux; a single oversized allocation, not gradual pressureOpen, no fix posted — headroom flags do not help a single allocation that is already too large
RAM/VRAM climbing only after upgrading#15759 — "Memory management changes and leaks after v0.30.2->v0.33.1 update"2026-08-20The newest report in this familyToo fresh for a verdict — read it before you take the v0.33.x upgrade
Allocation on device mid-sampleMany issues quote it—This is PyTorch's CUDA/HIP allocator string, not ComfyUI's--vram-headroom 2 first; --disable-dynamic-vram if that fails

Source: Comfy-Org/ComfyUI issue tracker, is:issue HostBuffer, read 2026-08-23. Run that search yourself before you act on any of it — states and comment counts move daily.


Why Does HostBuffer.read_file_slice Fail?

The single biggest ComfyUI memory story right now is issue #15255, and it is not a small-GPU problem. Its title, verbatim:

Dynamic VRAM streaming crashes all generations with HostBuffer.read_file_slice failed → CUDA OOM (regression after Aug 3 2026 update) (CORE-398)

Opened 2026-08-03, labelled Bug, assigned to maintainer rattus128, and still open on 2026-08-23. Sixty comments make it the most-commented open issue on the repository. The reason it deserves its own section rather than a row in a table: the word "regression" is doing real work. All generations crash, on setups that were fine the day before, which is a very different diagnosis from "your card ran out of room."

The signature is a pair of lines, in this order:

HostBuffer.read_file_slice failed
torch.AcceleratorError: CUDA error: out of memory

If you only read the second line you will conclude you need a bigger GPU. The first line is the one that identifies the bug — it comes from the comfy-aimdo streaming layer trying to read a slice of a model file off disk into a host buffer, and failing before anything about your VRAM capacity is relevant.

What the maintainer says. In a pinned comment dated 2026-08-06, rattus128 writes: "I have a working theory that this is a bug in Cuda." The thread notes it reproduces reliably only on multi-GPU configurations, and the maintainer has published a standalone diagnostic, windows_cuda_host_memory_diagnostic.py, asking affected users to run it from the same Python environment ComfyUI uses and post the output. PR #15610 is linked from the issue. Read the thread before you assume it is fixed — that status can change in a week.

What reporters in the thread found that worked:

# 1. The primary escape hatch — works for single- and multi-GPU alike
python main.py --disable-dynamic-vram

# 2. Multi-GPU only: pin ComfyUI to one device
python main.py --cuda-device 0

# 3. Multi-GPU only: if you must keep DynamicVRAM on
python main.py --disable-pinned-memory

Option 2 is the tell. If restricting ComfyUI to a single GPU makes an out-of-memory error disappear, the problem was never memory capacity — you just removed the multi-GPU host-buffer path the bug lives in. Anyone with two cards who has been buying a third to "fix" ComfyUI OOMs should try that one-line flag first; our best GPUs for image generation ranking is for people who genuinely need more silicon, not for people working around a streaming bug.


Should I Pin comfy-kitchen to 0.2.30?

There is a second HostBuffer failure with a completely different fix: a package downgrade, not a flag. Issue #15591, opened 2026-08-13 and still open, is titled:

[Bug] CUDA illegal memory access / HostBuffer.truncate failed with comfy-kitchen 0.2.31 during dynamic VRAM load (MiniMaxH3)

The two strings to match against your own console:

CUDA error: an illegal memory access was encountered
RuntimeError: HostBuffer.truncate failed

Note truncate, not read_file_slice — different call, different bug, different remedy. The reporter's crash happens during KSampler on a MiniMaxH3 video job (73 frames on a 12GB card), inside the partially_unload_ram / hostbuf.truncate step. The version boundary they identified is narrow and specific:

PackageVersionResult reported in the thread
comfy-kitchen0.2.31Crashes with HostBuffer.truncate failed
comfy-kitchen0.2.30Same workflow completes

So the workaround is one line, run inside the same virtual environment ComfyUI launches from:

pip show comfy-kitchen          # check what you actually have first
pip install comfy-kitchen==0.2.30

Two honest caveats. Pinning a dependency below what your ComfyUI build expects can break unrelated things, so treat this as a diagnostic probe rather than a permanent configuration — if 0.2.30 clears the crash, you have confirmed the cause and can wait for the fixed release instead of living on an old package. And this pin is specific to the truncate signature; it is not a general answer to the Aug 3 read_file_slice regression above, which reporters there address with flags.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

Why --lowvram Does Nothing Now

--lowvram was not removed. It was quietly demoted to a no-op on any default install.

Here is the exact line from comfy/cli_args.py:

vram_group.add_argument("--lowvram", action="store_true",
    help="Doesn't do anything if dynamic vram is enabled. If dynamic vram "
         "isn't being used this option makes the text encoders run on the CPU.")

The wording arrived in PR #13922, "Remove useless option and clarify what lowvram does", merged 2026-05-16. The behaviour change that made it necessary landed earlier:

DatePRWhat changed
2026-02-01#11845Adaptive model loading (AIMDO) added — "Reduce RAM usage, fix VRAM OOMs, and fix Windows shared memory spilling"
2026-02-28#12658"cli_args: Default comfy to DynamicVram mode" — this is the flip
2026-03-16#13002--enable-dynamic-vram added to force it on where it is off
2026-05-16#13922--lowvram help rewritten to admit it does nothing
2026-06-15#14480--reserve-vram reworked for DynamicVRAM; --vram-headroom added

Source: commit history of comfy/cli_args.py, Comfy-Org/ComfyUI, read 2026-08-18.

So the advice in every ComfyUI OOM thread predating March 2026 — including, until recently, the VRAM section of our own ComfyUI complete guide — is aimed at a memory manager you are no longer running.


Are You Even On It?

ComfyUI does not print "DynamicVRAM: enabled" at startup. You have to derive it from your own launch flags. This is the single most useful thing on this page, because it explains a contradiction people keep tripping over.

The gate is one function, reproduced verbatim from comfy/cli_args.py:

def enables_dynamic_vram():
    if args.enable_dynamic_vram:
        return True
    return (not args.disable_dynamic_vram
            and not args.highvram
            and not args.gpu_only
            and not args.novram
            and not args.cpu)

Read it as a table:

Your launch flagDynamicVRAM after that flag
(no flags at all)ON — this is the default
--disable-dynamic-vramOFF
--novramOFF (side effect)
--highvramOFF (side effect)
--gpu-onlyOFF (side effect)
--cpuOFF (side effect)
--lowvramstill ON — flag ignored
--enable-dynamic-vramON, forced, overrides everything above

That is the whole mystery. --novram, --highvram, --gpu-only and --cpu all turn DynamicVRAM off as a side effect of setting a different VRAM state — so the person who tried --novram in desperation reported "fixed!" and credited the wrong flag, while the person who followed the standard --lowvram advice saw literally no change and concluded the problem was their GPU. Both were looking at the same switch from opposite sides.

Two practical consequences. First, if someone in a thread tells you --novram fixed their OOM, try --disable-dynamic-vram instead — you probably get the fix without also forcing every model onto the CPU. Second, the startup line ComfyUI does print, Set vram state to: NORMAL_VRAM, tells you nothing about DynamicVRAM. Do not read it as confirmation either way.


Every Memory Flag, Verified

These are the memory-related arguments as they exist in comfy/cli_args.py on master at v0.33.1. Help text is quoted, not paraphrased.

FlagDefaultComfyUI's own help text
--disable-dynamic-vramoff"Disable dynamic VRAM and use estimate based model loading."
--enable-dynamic-vramoff"Enable dynamic VRAM on systems where it's not enabled by default."
--vram-headroom N0"Set the amount of vram in GB for DynamicVRAM to maintain as extra headroom above default. ComfyUI will try and keep this much VRAM completely free and unused, even counting VRAM from other apps."
--reserve-vram Nnone"Set the amount of vram in GB you want to reserve for use by your OS/other software. By default some amount is reserved depending on your OS."
--disable-nvml-pressureoff"Use CUDA instead of NVML for DynamicVRAM memory pressure."
--async-offload [N]on (NVIDIA), 2 streams"Use async weight offloading. An optional argument controls the amount of offload streams. Default is 2. Enabled by default on Nvidia."
--disable-async-offloadoff"Disable async weight offloading."
--fast-diskoff"Prefer disk-backed dynamic loading and offload over unpinned RAM. Can be faster for users with fast NVME disks."
--disable-pinned-memoryoff"Disable pinned memory use."
--disable-smart-memoryoff"Force ComfyUI to agressively offload to regular ram instead of keeping models in vram when it can." (typo is in the source)
--lowvramoff"Doesn't do anything if dynamic vram is enabled. If dynamic vram isn't being used this option makes the text encoders run on the CPU."
--novramoff"When lowvram isn't enough."
--highvramoff"By default models will be unloaded to CPU memory after being used. This option keeps them in GPU memory."
--gpu-onlyoff"Store and run everything (text encoders/CLIP models, etc... on the GPU)."
--cache-ram [GB [GB]]active 10% of RAM (min 2GB, max 10GB)"Use RAM pressure caching with the specified headroom thresholds. This is the default caching mode."
--cache-noneoff"Reduced RAM/VRAM usage at the expense of executing every node for each run."

--gpu-only, --highvram, --lowvram, --novram and --cpu are a mutually exclusive group in the parser — pass two and argparse rejects the launch. --disable-dynamic-vram is not in that group, so it composes freely with the others.

Two flags worth singling out. --vram-headroom is the one written for DynamicVRAM, and it is unusually aggressive: it keeps VRAM free "even counting VRAM from other apps", which is what you want on a machine where the same GPU is also driving your monitors. And --fast-disk is the opposite of a general-purpose fix — it prefers disk-backed offload, so it helps on fast NVMe and is exactly the wrong choice if your models live on a spinning disk or a network share. More on that in a moment.

If you are here because a specific model will not fit rather than because an update broke things, that is a different problem — see Run FLUX on a Low-VRAM GPU and FLUX VRAM requirements by GPU, or check the model against our VRAM calculator first.


Symptom → Cause → Fix

Ordered by how often each one is the real cause, judged by open-issue volume on Comfy-Org/ComfyUI as of 2026-08-23. Every issue number below is a real, open report you can read yourself. The two HostBuffer signatures have their own sections above (#15255 and #15591) because their fixes differ from everything here.

1. Everything got slower after an update, and your models are on a hard drive or network share

The most-reported flavour of this regression. DynamicVRAM streams weights on demand instead of loading them once, so when the backing store is slow, it re-reads constantly. Issue #15661 (opened 2026-08-16) is titled "Performance regression after recent updates: DynamicVRAM / AIMDO causes extreme slowdowns when models are on HDD" and attaches a feature request for a "prefer GPU / disable dynamic memory" option — which tells you the reporter had not found --disable-dynamic-vram either.

Fix: --disable-dynamic-vram. Do not reach for --fast-disk here; it prefers disk-backed loading, which is the direction you are trying to run away from. --fast-disk is for NVMe owners who want more disk streaming, not less.

2. Straight OOM: "Allocation on device"

The phrase people paste into Google is the exact torch allocator string Allocation on device, and the important thing to know is that it comes from PyTorch's CUDA/HIP allocator, not from ComfyUI — so it tells you memory ran out, and nothing at all about why. Under DynamicVRAM it can now fire mid-sample rather than at load time, which is why it feels different from the OOMs you are used to. You can see the current volume for yourself with the tracker's own search: is:issue "out of memory".

Before you treat this as a capacity problem, scroll up one line in your console. If HostBuffer.read_file_slice failed precedes it, you are in the Aug 3 regression, not out of VRAM. Issue #14340 (2026-06-08) is the other flavour worth knowing — "VRAM OOM on Linux with large singular allocation but should be within limits" — where a single oversized allocation fails even though total free memory looks adequate.

Fix, in order: try --disable-dynamic-vram first. If you want to keep DynamicVRAM, give it room instead: --vram-headroom 2 (start at 2GB) so it stops competing with your desktop compositor and browser for the last gigabyte.

3. Infinite hang at "Model Initializing..."

Not a crash, not an error — the UI just stops. Issue #15628 (opened 2026-08-14) reports it in its title: "[Bug] MiniMaxH3: DynamicVRAM causes infinite hang at 'Model Initializing...' on RTX 4070 12GB; custom nodes surface underlying CUDA illegal memory access." Its close relative is #15591 (2026-08-13), "[Bug] CUDA illegal memory access / HostBuffer.truncate failed with comfy-kitchen 0.2.31 during dynamic VRAM load (MiniMaxH3)" — same model family, same illegal-memory-access root, but that one has a package-level fix rather than a flag-level one.

Fix: --disable-dynamic-vram. If you see HostBuffer.truncate failed anywhere in the traceback, go to the comfy-kitchen version pin instead — flags will not clear that one. If the hang survives both, it is not this bug and you should check your custom nodes; a node throwing during load looks identical from the front end.

4. WSL: model set pinned into "Shared GPU memory" for no reason

Issue #15679 (opened 2026-08-16): "DynamicVRAM on WSL pins the entire model set into Shared GPU memory for no benefit." Shared GPU memory is system RAM presented to the GPU through the WSL driver, and it is dramatically slower than real VRAM. Windows Task Manager shows this clearly — if your dedicated VRAM is half empty while shared GPU memory is pegged, you are in it.

Fix: --disable-dynamic-vram, and consider --disable-pinned-memory if the shared-memory usage persists.

5. --reserve-vram appears to do nothing

Issue #15666 (opened 2026-08-16) is literally titled "--reserve-vram ignored", and it was still open on 2026-08-23. The flag was reworked for DynamicVRAM in PR #14480 (2026-06-15). Argparse accepts it and the launch succeeds, which is exactly why it wastes people's time — there is no error to tell you the value was discarded.

Fix: use --vram-headroom N instead — it is the flag actually written for the DynamicVRAM path. If you need the old semantics, --disable-dynamic-vram restores the estimate-based loader that --reserve-vram was designed against.

6. AMD / ROCm: blank or corrupted outputs

Issue #15436 (opened 2026-08-08): "Blank invalid/outputs using dynamic vram on ROCM 7.14 on gfx1201". Issue #15484 (2026-08-11) reports a ROCm/Windows VAE slowdown under Dynamic VRAM that a forced unload stabilises. Note the symptom here is wrong pixels, not an error — you can lose an afternoon before you suspect memory management at all.

Fix: --disable-dynamic-vram before you start bisecting your workflow. If your ComfyUI is producing black or empty images generally, that has several unrelated causes too.

7. Host buffer grows until RAM is exhausted on long runs

Issue #15575 (2026-08-13) documents an AIMDO host buffer growing across a single long prompt — "618x hostbuf_grow ERROR, ~12% slowdown" with a LoRA-loaded workflow. If your system RAM climbs steadily over a long batch and never comes back down, this is a candidate.

Fix: --disable-dynamic-vram, or cap the RAM cache with --cache-ram 4 and see whether the growth stops.

8. First generation is fine, the second one dies

The most misleading pattern of the lot, because a successful first run convinces you your setup is correct. Issue #15269 (opened 2026-08-03) is titled "DynamicVRAM + AIMDO 0.4.11 causes 'VRAM Allocation failed (non OOM)' and 'Fault failed: 2' on second generation (QwenImage)". Note the words non OOM in ComfyUI's own message: the allocator is reporting a failure that is explicitly not an out-of-memory condition, which is the manager admitting the fault is its own. The same class of failure has been open since March in issue #12943, "aimdo:VRAM Allocation failed (non OOM) in 17.1", so this is not purely an August problem.

Fix: --disable-dynamic-vram. Neither issue has a maintainer fix as of 2026-08-23, so the flag is the only lever on your side of the machine. If you are on QwenImage specifically, keep an eye on #15269 rather than rebuilding your workflow.


What This Page Does Not Claim

We do not own the hardware to benchmark this, so there is no invented performance table anywhere on this page. That matters, because the honest state of DynamicVRAM is that it is a genuine improvement on some configurations and a serious regression on others, and nobody — including NVIDIA, Comfy-Org, or us — has published the split.

What is source-verified: the flags exist, the defaults are what we describe, --lowvram is inert, and every issue number, title and error string above was read off the public tracker on 2026-08-23. What we are explicitly not claiming: that --disable-dynamic-vram will be faster for you. On a machine with fast NVMe storage and plenty of system RAM, DynamicVRAM is doing useful work and turning it off may cost you the ability to run a model at all.

We are also not claiming a root cause for the Aug 3 regression. The maintainer's own pinned position is a theory — "I have a working theory that this is a bug in Cuda" — and a theory from the person who wrote the code is still a theory. The pattern in the issue tracker is that the losers are consistently multi-GPU rigs, HDD or network-backed model storage, WSL, and ROCm. If none of those describe you, benchmark before you disable anything permanently.


Measure It Yourself

Two commands settle it in five minutes. First, confirm which flags your install actually has, rather than trusting any article:

python main.py --help | grep -A2 -- "--disable-dynamic-vram"
python main.py --help | grep -A2 -- "--lowvram"

If the --lowvram help text on your machine still mentions dynamic vram, you are on a build that matches this page. If those flags do not exist at all, you are on something older than 2026-03-16 and none of this applies to you.

Then run the same workflow twice, changing exactly one thing:

# Run A — default (DynamicVRAM on)
python main.py

# Run B — same workflow, manager off
python main.py --disable-dynamic-vram

Watch VRAM live in a second terminal with nvidia-smi --query-gpu=memory.used --format=csv -l 1 (or rocm-smi on AMD), and take wall-clock time from the ComfyUI console, which prints "Prompt executed in X seconds" after every run. Record peak VRAM, wall-clock, and whether it OOMs. Three numbers, two runs, and you know the answer for your machine instead of somebody else's.

Do this with a workflow you actually use. An SDXL graph and a video graph can land on opposite sides of this — video models are the ones filling the issue tracker, and a clean SDXL result tells you nothing about them. If you are setting up a comparison from scratch, our ComfyUI FLUX workflow guide and SDXL vs FLUX comparison both include graphs you can reuse as a fixed baseline.


Verdict

  1. If it broke after an update, try --disable-dynamic-vram first. One flag, instantly reversible, and it addresses seven of the eight symptoms above.
  2. Read the line above the OOM. HostBuffer.read_file_slice failed means #15255 and a flag; HostBuffer.truncate failed means #15591 and pip install comfy-kitchen==0.2.30. Same word, different bug, different fix — and neither one means your GPU is too small.
  3. Stop adding --lowvram. It is a no-op on a default install and has been since the February 2026 default flip. The source says so in plain English.
  4. --novram "working" is a red herring. It disables DynamicVRAM as a side effect while also forcing everything to CPU. You almost certainly want just the first half.
  5. On two or more GPUs, try --cuda-device 0 before you buy anything. If pinning to a single card clears the OOM, you did not have a capacity problem — you had the multi-GPU host-buffer bug.
  6. If you want to keep DynamicVRAM, give it headroom. --vram-headroom 2 is a gentler intervention than switching the whole manager off, and it accounts for VRAM used by other applications.
  7. Benchmark before you make it permanent. DynamicVRAM is not a mistake — it is a real improvement on the right hardware. Multi-GPU rigs, HDD-backed models, WSL and ROCm are where it currently hurts.

Recheck date: mid-November 2026. ComfyUI shipped v0.24.0 through v0.33.1 between 2026-06-03 and 2026-08-13 — twelve tagged releases in ten weeks — and every issue cited here was still open on 2026-08-23. #15255 in particular has an assignee and a linked PR, so it may well be resolved before you read this; check the issue before you apply the workaround. Flag names and defaults in this area have a short shelf life, and the --help command above is always more current than we are.

For hardware planning rather than firefighting, our best GPUs for image generation ranking covers what actually clears these workloads without flag gymnastics.


Sources

  • comfy/cli_args.py, ComfyUI master — every flag name, default and help string quoted above; read 2026-08-18
  • Issue #15255 — "Dynamic VRAM streaming crashes all generations with HostBuffer.read_file_slice failed → CUDA OOM (regression after Aug 3 2026 update) (CORE-398)"; opened 2026-08-03, open with 60 comments, assigned to rattus128, pinned maintainer comment 2026-08-06, linked PR #15610
  • Issue #15591 — "[Bug] CUDA illegal memory access / HostBuffer.truncate failed with comfy-kitchen 0.2.31 during dynamic VRAM load (MiniMaxH3)"; opened 2026-08-13; source of the 0.2.31 broken / 0.2.30 working boundary
  • Issue #15269 — "DynamicVRAM + AIMDO 0.4.11 causes 'VRAM Allocation failed (non OOM)' and 'Fault failed: 2' on second generation (QwenImage)"; opened 2026-08-03
  • Issue #14340 — "VRAM OOM on Linux with large singular allocation but should be within limits"; opened 2026-06-08
  • Comfy-Org/ComfyUI issue tracker — issues #15759, #15679, #15666, #15661, #15628, #15575, #15484, #15436, #12943 — all confirmed open with the titles quoted above on 2026-08-23
  • ComfyUI releases — v0.33.1 published 2026-08-13; release cadence v0.24.0 (2026-06-03) through v0.33.1
  • Commit history of comfy/cli_args.py — PRs #11845, #12658, #13002, #13922, #14480 with merge dates

FAQ

🎯
AI Learning Path

Generating images locally? Take it further.

From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Once your hardware is sorted

Go from one-off images to a real workflow

The Local Image Generation course covers ComfyUI, SDXL and FLUX properly — plus 24 more courses on running AI on your own hardware.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsComfyUIDynamicVRAMAIMDOHostBufferOut of MemoryVRAMTroubleshooting

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on Local Image Generation
See the full Run FLUX.1 Locally guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

What does "HostBuffer.read_file_slice failed" mean in ComfyUI?

It means the comfy-aimdo dynamic VRAM streaming layer failed while reading a slice of a model file into a host buffer — it is not a statement about your GPU capacity, even though a "torch.AcceleratorError: CUDA error: out of memory" traceback usually follows it one line later. This is tracked as issue #15255, "Dynamic VRAM streaming crashes all generations with HostBuffer.read_file_slice failed → CUDA OOM (regression after Aug 3 2026 update) (CORE-398)", opened 2026-08-03 and still open on 2026-08-23 with 60 comments — the most-commented open issue on the repository. Reporters clear it with --disable-dynamic-vram; on multi-GPU machines --cuda-device 0 or --disable-pinned-memory also worked.

Is HostBuffer.truncate failed the same bug as read_file_slice failed?

No, and the fix is different, which is why the distinction is worth a minute of your time. "RuntimeError: HostBuffer.truncate failed", usually paired with "CUDA error: an illegal memory access was encountered", is issue #15591 — "[Bug] CUDA illegal memory access / HostBuffer.truncate failed with comfy-kitchen 0.2.31 during dynamic VRAM load (MiniMaxH3)", opened 2026-08-13. The reporter narrowed it to a single package version: comfy-kitchen 0.2.31 crashes, 0.2.30 completes the same workflow. So the workaround there is pip install comfy-kitchen==0.2.30 inside your ComfyUI virtual environment, not a launch flag. Treat the downgrade as a diagnostic probe rather than a permanent setup.

Why does ComfyUI say "VRAM Allocation failed (non OOM)"?

Read the parenthesis literally: the memory manager is telling you the allocation failed for a reason that is not running out of memory. It shows up in issue #15269, "DynamicVRAM + AIMDO 0.4.11 causes 'VRAM Allocation failed (non OOM)' and 'Fault failed: 2' on second generation (QwenImage)" (opened 2026-08-03), where the first generation succeeds and the second one dies — and in the much older #12943, "aimdo:VRAM Allocation failed (non OOM) in 17.1", open since 2026-03-14. Neither has a maintainer fix as of 2026-08-23, so --disable-dynamic-vram is the only lever available to you. Buying a bigger card will not help a non-OOM allocation failure.

Why did --lowvram stop working in ComfyUI?

Because it is now inert whenever DynamicVRAM is active, which is the default. ComfyUI's own help text for the flag, in comfy/cli_args.py, reads verbatim: "Doesn't do anything if dynamic vram is enabled. If dynamic vram isn't being used this option makes the text encoders run on the CPU." The clarification landed in PR #13922 ("Remove useless option and clarify what lowvram does") on 2026-05-16. Every tutorial and forum reply telling you to add --lowvram was written before that, and adding it changes nothing on a default install.

What is DynamicVRAM in ComfyUI?

It is the adaptive model-loading and offloading layer built on the comfy_aimdo package, introduced in PR #11845 on 2026-02-01 and made the default mode by PR #12658 ("cli_args: Default comfy to DynamicVram mode") on 2026-02-28. Instead of estimating how much VRAM a model needs before loading it, DynamicVRAM loads and evicts weights on demand while sampling runs, using NVML memory pressure readings to decide. When it works, it fits models that used to OOM. When it misjudges your storage or your platform, it thrashes.

How do I turn DynamicVRAM off?

Add --disable-dynamic-vram to your launch command and restart ComfyUI. Its help text is "Disable dynamic VRAM and use estimate based model loading" — that is the pre-2026 behaviour every older guide assumes. Note that --novram, --highvram, --gpu-only and --cpu also disable it as a side effect, which is why people who tried --novram reported their problem "fixed" while people who tried --lowvram saw no change at all.

Is DynamicVRAM enabled by default on AMD and ROCm?

The gating function enables_dynamic_vram() in comfy/cli_args.py contains no platform check on current master — it returns true unless you pass --disable-dynamic-vram, --highvram, --gpu-only, --novram or --cpu. ROCm users are filing DynamicVRAM-specific bugs (issue #15436, "Blank invalid/outputs using dynamic vram on ROCM 7.14 on gfx1201", opened 2026-08-08), so it is clearly reaching them. What is not visible from the source alone is whether the comfy_aimdo backend actually engages on every build, and ComfyUI does not print a "DynamicVRAM: on" line at startup. Do not assume — test with and without the flag on your own machine.

Why is --reserve-vram being ignored?

That is an open, unresolved bug at the time of writing: issue #15666, "--reserve-vram ignored", opened 2026-08-16 against Comfy-Org/ComfyUI. The flag was reworked for DynamicVRAM in PR #14480 on 2026-06-15, which is also when --vram-headroom was added. If you need ComfyUI to leave VRAM free for your display or another app right now, --vram-headroom is the flag written specifically for DynamicVRAM: "Set the amount of vram in GB for DynamicVRAM to maintain as extra headroom above default."

Should I just roll back to an older ComfyUI version?

You would have to go back a long way. DynamicVRAM became the default on 2026-02-28, so any release from spring 2026 onward has it. Between 2026-06-03 and 2026-08-13 alone ComfyUI shipped v0.24.0 through v0.33.1 — twelve tagged releases in ten weeks — and rolling back that far costs you every model and node added since. Flipping one flag is cheaper and reversible. The one exception worth knowing is a package-level pin rather than an app-level rollback: if your traceback contains HostBuffer.truncate failed, issue #15591 identifies comfy-kitchen 0.2.31 as the bad version and 0.2.30 as the working one, which is a far smaller step backwards than downgrading ComfyUI itself. If you are considering the v0.33.x upgrade, issue #15759 ("Memory management changes and leaks after v0.30.2->v0.33.1 update", opened 2026-08-20) is worth reading first.

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks for local models — reproducible environments that survive an update. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: September 13, 2026🔄 Last Updated: September 13, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Generating images locally? Take it further.

From FLUX and ComfyUI setup to building real image pipelines and apps. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators