Qwen3.8-27B VRAM Requirements: What Runs on 16GB, 24GB and 32GB
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Got the hardware sorted? Now build on it. You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.
Short answer: Qwen3.8-27B wants a 24GB card. Ollama ships one build, qwen3.8:27b, and it is 18GB before you add a single token of context, so an RTX 3090, 4090 or 7900 XTX runs it at 4-bit with room for a 64K context. A 16GB card (RTX 4080, 5070 Ti, 5080, 4060 Ti 16GB) runs it only at 3-bit, from a 13.1 GB GGUF file you load in LM Studio or llama.cpp rather than Ollama. A 32GB card (RTX 5090) is where the model stops being squeezed: Q6_K with the full 128K context, or Q8_0 with 48K. Below 16GB, run a different model. Everything on this page is arithmetic from the published files and the model's own config, and every speed figure is a bandwidth ceiling, not a measurement.
Qwen released Qwen3.8 in exactly two sizes: this dense 27B and a 2.4-trillion-parameter MoE that is not a consumer conversation. There is no 8B, no 14B. So the entire question of running Qwen3.8 at home collapses into one number: does 27 billion parameters fit the card you have, and at what precision. This page answers that tier by tier, with file sizes read off the Unsloth GGUF repository and Ollama's tag list, and the KV cache worked out from the model's actual attention layout rather than a rule of thumb.
Fit table by VRAM tier
24GB is the tier the model was built for. 16GB works at 3-bit. 32GB buys precision and context, not a different model.
Cards are quoted in GiB because that is how VRAM is sold. File sizes are the published decimal-GB figures from unsloth/Qwen3.8-27B-GGUF, converted to GiB so they compare against your card. "Context you can afford" assumes about 1 GiB of runtime overhead and f16 KV cache at 64 KiB per token (worked out in the KV cache section).
| Card class | Best quant | Weights on disk | Context you can afford | Verdict |
|---|---|---|---|---|
| 8GB (RTX 4060, 3060 Ti, 5060) | UD-IQ1_M loads, do not bother | 6.73 GB / 6.3 GiB | ~8K | Not a real option. 1-bit 27B loses more than an 8B model at Q4 ever does. |
| 12GB (RTX 3060 12GB, 4070, 5070) | UD-Q2_K_XL | 9.83 GB / 9.2 GiB | ~16K | Loads at 2-bit. Quality cost is steep; a 14B at Q4 is the better 12GB pick. |
| 16GB (RTX 4080, 5070 Ti, 5080, 4060 Ti 16GB, 5060 Ti 16GB) | UD-Q3_K_XL (GGUF, not Ollama) | 13.1 GB / 12.2 GiB | ~32K | Works. The biggest Qwen a 16GB card runs, at 3-bit, via LM Studio or llama.cpp. |
| 24GB (RTX 3090, 4090, RX 7900 XTX) | Q4_K_M (ollama pull qwen3.8:27b) | 17 GB + 0.93 GB projector / ~16.7 GiB | ~64K | The intended home. Q5_K_M (19.8 GB) also fits with ~48K; Q6_K (22 GB) fits with ~32K and no projector. |
| 32GB (RTX 5090, Radeon AI Pro R9700) | Q6_K for context, Q8_0 for precision | 22 GB / 20.5 GiB, or 29 GB / 27 GiB | ~128K at Q6_K, ~48K at Q8_0 | Near-lossless. Ollama has qwen3.8:27b-q8_0 at 30GB. |
| 48GB (RTX 6000 Ada, A6000, 2 x 24GB) | Q8_0 | 29 GB / 27 GiB | Full 262K (16 GiB) | Everything on one card. BF16 at 56GB still does not fit. |
| Apple 24GB unified (M4 Pro) | UD-Q3_K_XL | 13.1 GB / 12.2 GiB | ~16K | macOS keeps several GB for itself; treat 24GB as roughly 17GB usable. The 18GB MLX tag is too tight here. |
| Apple 48GB+ unified (M4 Max) | qwen3.8:27b-mlx or Q6_K | 18GB, or 22 GB | ~128K | Comfortable. 64GB and up runs Q8_0 with the full context. |
Two things to check before you trust any row against your own card. First, ollama ps after loading should say 100% GPU; any CPU share means you spilled, and on this model the culprit is almost always context rather than weights. Second, the vision projector is optional: 931MB in Ollama's tag, 928MB as mmproj-F16.gguf in the Unsloth repo. Leave it out on a 16GB card unless you need image input, and that memory becomes context.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
Where 18GB comes from
Ollama's 18GB is a 17GB Q4_K_M model plus a 931MB vision projector. Parameter count times bytes per weight predicts it within 3%.
The footprint method used across this site is the one you can audit: parameter count times llama.cpp's published bits per weight for each k-quant. For a 27B model:
| Quant | Bytes per weight (GB per billion) | 27B footprint, computed | Published Unsloth file | Ollama tag |
|---|---|---|---|---|
| Q3_K_S class | ~0.43 | ~11.6 GB | UD-Q3_K_XL 13.1 GB (dynamic quant keeps more layers at higher bits) | none |
| Q4_K_M | ~0.60 | ~16.2 GB | UD-Q4_K_M 16.5 GB | qwen3.8:27b, 17GB + projector = 18GB |
| Q5_K_M | ~0.71 | ~19.2 GB | UD-Q5_K_M 19.8 GB | none |
| Q6_K | ~0.82 | ~22.1 GB | UD-Q6_K 22 GB | none |
| Q8_0 | ~1.06 | ~28.6 GB | Q8_0 29 GB | qwen3.8:27b-q8_0, 30GB |
| BF16 | 2.0 | ~54 GB | (split files) | qwen3.8:27b-bf16, 56GB |
The Ollama tag page lists the model at 27.3B parameters, which is why the published Q4 file lands a little above the 27 x 0.60 estimate. The UD- prefix on the Unsloth files marks their dynamic quants, which keep sensitive layers at higher precision instead of flattening everything to one bit-width; that is why UD-Q3_K_XL is heavier than a plain Q3_K_S estimate and also why it is the 3-bit file worth using.
Ollama's twelve tags are all this one 27B model in different containers: 27b / 27b-q4_K_M / 27b-nvfp4 / 27b-mlx / 27b-mtp-q4_K_M at 18GB, 27b-q8_0 / 27b-mtp-q8_0 at 30GB, 27b-mxfp8 at 32GB, and the three BF16 variants at 56GB. Anything smaller than 18GB is not in Ollama's registry at all, which is the single most important fact for 16GB owners.
The KV cache bill
64 KiB per token at f16. That is 2 GiB at 32K, 4 GiB at 64K, 8 GiB at 128K, 16 GiB at the full 262,144.
This is the part most VRAM pages get wrong for Qwen3.8, because they apply a conventional-transformer rule of thumb to a model that is not one. Qwen3.8-27B is a hybrid. Its model card gives the layout as 16 blocks of (3 x Gated DeltaNet then 1 x Gated Attention), and the config.json confirms it: 64 layers, full_attention_interval: 4, 4 KV heads, head dimension 256.
The 48 Gated DeltaNet layers are linear attention. They carry a fixed-size recurrent state that does not grow with the prompt, a few tens of megabytes in total. Only the 16 full-attention layers keep a cache that scales with context, and the arithmetic for those is:
K and V x 4 KV heads x 256 dims x 2 bytes (f16) = 4 KiB per token per layer
4 KiB x 16 full-attention layers = 64 KiB per token
8K context -> 0.5 GiB
32K context -> 2 GiB
64K context -> 4 GiB
128K context -> 8 GiB
262K context -> 16 GiB
For comparison, a 64-layer model where every layer had the same 4 x 256 attention heads would cost 256 KiB per token, or 8 GiB at 32K. The hybrid layout cuts that bill by four, and it is the reason a 24GB card can hold an 18GB build and still afford a 64K window. Two practical notes:
- Ollama's default context is not the model's. Ollama sets a modest
num_ctxunless you ask for more. A 27B model with a 262K native window and a 4K working context is wasting the card. Set it explicitly (/set parameter num_ctx 65536in the REPL, or in a Modelfile). - Thinking eats context. Qwen3.8 turns thinking on by default at
reasoning_effort="xhigh", and Unsloth's recommended agentic budget is 262,144 reasoning tokens with 131,072 for the answer. On a card that is nearly full, a long thinking trace is exactly what pushes the KV cache over the edge. Turn the effort down tomediumorlowfor chat, or disable thinking per request. - KV quantisation is the escape hatch. llama.cpp and LM Studio can store the cache at q8_0, halving the per-token cost to 32 KiB. On 16GB, that turns a ~32K budget into ~64K.
16GB: the 3-bit route
UD-Q3_K_XL at 13.1 GB (12.2 GiB) is the file. It is the biggest Qwen a 16GB card runs, and you have to leave Ollama to get it.
The maths on a 16 GiB card: 16 minus 12.2 for weights minus roughly 1 for the CUDA context and compute buffers leaves about 2.5 to 2.8 GiB. At 64 KiB per token that is a 32K context with a little to spare, or 64K with q8_0 KV cache. That is a genuinely usable configuration for chat and single-file coding, and a reasonable one for a coding agent if you set num_ctx deliberately.
What does not work at 16GB, despite looking close:
- UD-IQ4_XS, 14.3 GB (13.3 GiB). Fits, with about 1.5 GiB left after overhead. That is a 16K to 24K context, no vision projector. Fine for short prompts, wrong for agents.
- UD-Q4_K_S, 15.4 GB (14.3 GiB). Loads with under a gigabyte free. An 8K context at best. Not worth the fight.
- Ollama's
qwen3.8:27bat 18GB. Spills into system RAM before the first token and speed collapses. Ollama does not publish anything smaller.
The route, using the Hugging Face CLI and then either LM Studio's model folder or llama.cpp directly:
huggingface-cli download unsloth/Qwen3.8-27B-GGUF --include "*UD-Q3_K_XL*" --local-dir ./qwen3.8-27b
# optional, only if you want image input (+0.87 GiB):
huggingface-cli download unsloth/Qwen3.8-27B-GGUF --include "mmproj-F16.gguf" --local-dir ./qwen3.8-27b
# llama.cpp, 32K context, all layers on the GPU
llama-server -m ./qwen3.8-27b/Qwen3.8-27B-UD-Q3_K_XL.gguf -c 32768 -ngl 99 --temp 1.0 --top-p 0.95 --top-k 20
The sampling values are Unsloth's recommended thinking-mode settings for this model (temperature 1.0, top_p 0.95, top_k 20, min_p 0). For non-thinking use they suggest temperature 0.7, top_p 0.8 and a presence penalty of 1.5.
Whether a 3-bit 27B beats a Q4 14B on your work is not something anyone has published a clean head-to-head on, and we are not going to invent one. The Qwen 16GB page lays out the trade: the 27B knows more and reasons harder; the 14B keeps full-fidelity weights and roughly twice the context headroom. Pull both and run your own prompts.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
24GB: the intended home
One command, 18GB, and a 64K context left over. This is the tier Qwen3.8-27B was tuned for.
On a 24 GiB card the default Ollama build (17 GB model plus 931MB projector, about 16.7 GiB) leaves 24 minus 16.7 minus 1 of overhead, roughly 6 GiB, which is a 96K f16 context on paper and a comfortable 64K in practice once you leave a margin for prompt-processing buffers. Drop the vision projector and you gain almost another 15K.
Your other options at 24GB:
- Q5_K_M, 19.8 GB (18.4 GiB). About 4.5 GiB left, a 64K context with q8_0 KV cache or 48K at f16. The quality step from Q4 to Q5 is small but real on long-form and code.
- Q6_K, 22 GB (20.5 GiB). About 2.5 GiB left, so a 32K context and no vision projector. Near-lossless weights, cramped everything else. On 24GB this is a chat configuration, not an agent one.
- UD-Q4_K_XL, 17.6 GB (16.4 GiB). Unsloth's dynamic 4-bit, slightly heavier than Q4_K_M with a few more layers kept at higher precision. The Ollama default is the simpler choice unless you are already loading GGUFs by hand.
If you own a 24GB card and were on Qwen3.6-27B, the swap is like-for-like on memory: Ollama lists qwen3.6:27b at 17GB and qwen3.8:27b at 18GB. What you gain is native image and video input, a 262K native context in place of 128K, and the benchmark gains Qwen reports on its model card. What you lose is the non-thinking default; see the model page for the full upgrade verdict.
32GB and up: Q6, Q8 and long context
32GB does not change which Qwen you run. It changes how many bits each weight keeps and how much of the 262K window you can actually use.
On an RTX 5090 or a 32GB Radeon AI Pro R9700:
- Q6_K, 22 GB (20.5 GiB), for context. About 10.5 GiB free after overhead: the full 128K context at f16 (8 GiB) with the projector loaded. This is the configuration for long documents and repo-wide agents.
- Q8_0, 29 GB (27 GiB), for precision. About 4 GiB free, a 48K to 64K context depending on KV quantisation. Ollama ships it as
qwen3.8:27b-q8_0at 30GB, so it is a one-liner. Effectively lossless weights. qwen3.8:27b-mxfp8, 32GB. Fills the card with weights alone. Skip it on a 32GB GPU.
At 48GB (RTX 6000 Ada, A6000, or two 24GB cards with layer split) Q8_0 plus the entire 262,144-token context (16 GiB) fits on the GPU with room to spare, which is the first tier where the model's native window is actually usable end to end. The BF16 build is 56GB and still does not fit 48GB; that one is for 64GB unified-memory Macs and workstation cards.
8GB and 12GB: run something else
The files exist. The quality does not. Below 16GB, a smaller model at 4-bit beats this 27B at 1 or 2 bits.
Unsloth publishes builds all the way down to UD-IQ1_S at 6.19 GB, so it is technically true that an 8GB card can load Qwen3.8-27B. It is also true that a 1-bit or 2-bit quant of a 27B has thrown away more information than a good 8B or 14B at Q4 ever had to. The Qwen 8GB picks and 12GB VRAM picks list the models built for those cards.
The one defensible 12GB configuration is partial offload: run the 18GB Ollama build and let Ollama put the layers that do not fit into system RAM. It works, it keeps 4-bit quality, and it runs at a few tokens per second because every token crosses the PCIe bus. If you have 32GB of system RAM and patience, that beats 2-bit. If you want it to feel fast, you want a bigger card.
Speed ceilings by card
Ceilings, not measurements. Nobody here owns these cards, so every figure is memory bandwidth divided by the gigabytes a token has to read.
Token generation on a dense model reads the whole weight set once per token, so the hard upper bound is bandwidth over footprint. Real engines land materially below it: attention, the DeltaNet recurrence, sampling and kernel launches all cost time the ceiling ignores, and a hybrid model's linear-attention layers add compute that plain bandwidth does not capture. Read each figure as the speed the card cannot exceed.
| Card | VRAM | Bandwidth | Quant it runs | GB read per token | Ceiling (tok/s) |
|---|---|---|---|---|---|
| RTX 5090 | 32GB | 1,792 GB/s (512-bit x 28 Gbps) | Q8_0 / Q4_K_M | 29 / 16.5 | ~62 / ~109 |
| RTX 4090 | 24GB | 1,008 GB/s | Q4_K_M | 16.5 | ~61 |
| RTX 3090 | 24GB | 936 GB/s | Q4_K_M | 16.5 | ~57 |
| RTX 5080 | 16GB | 960 GB/s | UD-Q3_K_XL | 13.1 | ~73 |
| RTX 5070 Ti | 16GB | 896 GB/s (256-bit x 28 Gbps) | UD-Q3_K_XL | 13.1 | ~68 |
| RTX 4080 | 16GB | 717 GB/s (256-bit x 22.4 Gbps) | UD-Q3_K_XL | 13.1 | ~55 |
| RTX 5060 Ti 16GB | 16GB | 448 GB/s (128-bit x 28 Gbps) | UD-Q3_K_XL | 13.1 | ~34 |
| RTX 4060 Ti 16GB | 16GB | 288 GB/s (128-bit x 18 Gbps) | UD-Q3_K_XL | 13.1 | ~22 |
Two things fall out of that table. The 16GB tier is bandwidth-split down the middle: a 5080 or 5070 Ti has a higher ceiling on the 3-bit file than a 4090 has on the 4-bit one, while a 4060 Ti 16GB is a fifth of that. And the 5090's advantage at Q4 is mostly wasted; it earns its price at Q8 and at 128K context, where a 24GB card cannot follow.
Ollama's -mtp tags bundle the model's multi-token-prediction head, which Qwen trained alongside the main weights and Unsloth ships as a separate 1.37 GB Q4_0 file. It exists to draft several tokens per step for speculative-style decoding. Whether it lifts your card's throughput, and by how much, is something to measure on your own machine; we have no figure and will not guess one.
Ollama pull commands
One line per tier. The 16GB route is not an Ollama command, because Ollama has no sub-18GB build.
# 24GB cards (RTX 3090 / 4090 / 7900 XTX): the default 4-bit build, 18GB
ollama pull qwen3.8:27b
ollama run qwen3.8:27b
# then, inside the REPL, give it a real context window:
# /set parameter num_ctx 65536
# 32GB cards (RTX 5090): 8-bit, 30GB
ollama pull qwen3.8:27b-q8_0
# Apple Silicon with 48GB+ unified memory
ollama pull qwen3.8:27b-mlx
# 16GB cards: not Ollama. See the 3-bit route above.
huggingface-cli download unsloth/Qwen3.8-27B-GGUF --include "*UD-Q3_K_XL*" --local-dir ./qwen3.8-27b
To turn thinking down for everyday chat, pass reasoning_effort in the request (low, medium or the default xhigh), or set it in a Modelfile so every session inherits it. For a longer walk through Ollama parameters on the Qwen family, the Qwen 3 local setup guide covers Modelfiles, num_ctx and flash attention.
What to buy if you do not have it
If you are buying a card for this model, buy 24GB. 16GB is a compromise you will feel at 3-bit; 32GB is a luxury you will feel only at Q8 and long context.
- Already on 16GB and happy at 3-bit? Stay. Read the best LLM for 16GB VRAM page for the models that fit that card without compromise, and keep Qwen3.8-27B as the heavy hitter you load for hard problems.
- Buying 16GB new? The RTX 5080 and 5070 Ti are the bandwidth-rich 16GB cards, and what runs on an RTX 5080 covers exactly which models that card handles at full quality. The 4060 Ti 16GB has the memory but a quarter of the bandwidth; every dense model on it runs at a quarter of the speed.
- Buying for this model specifically? A used RTX 3090 is the cheapest 24GB with 936 GB/s, and it runs the 18GB build with a 64K context. The best LLM for 24GB VRAM page shows what else that tier unlocks, which is most of the interesting open-weight catalogue.
- Want the full 128K context and Q8? That is a 32GB RTX 5090 or a 48GB workstation card, and the hardware hub compares them against the unified-memory route, where a 64GB Mac or a 128GB Strix Halo box runs Q8 with the whole 262K window and no card at all.
Whichever tier you land in, check the VRAM calculator with your exact card and context before you pull 18GB over a slow connection.
Sources
- Qwen/Qwen3.8-27B model card: dense 27B, 64 layers, 16 x (3 x Gated DeltaNet then 1 x Gated Attention), 24 Q / 4 KV heads, 262,144-token native context extensible to 1M, Apache 2.0, image and video input, thinking on by default at xhigh.
- Qwen/Qwen3.8-27B config.json: hidden_size 5120, num_hidden_layers 64, num_key_value_heads 4, head_dim 256, full_attention_interval 4, max_position_embeddings 262144.
- Ollama library, qwen3.8 tags: twelve tags, all 27B and 256K context; the default 18GB tag is a 17GB Q4_K_M model plus a 931MB projector, listed at 27.3B parameters.
- unsloth/Qwen3.8-27B-GGUF file list and model card: every file size quoted above, the recommended sampling settings, and the MTP note.
- Footprint method: llama.cpp's published bits per weight for each k-quant, the same table used on the RTX 5080 and RTX 5090 pages. Bandwidths are NVIDIA's bus width times memory speed, with the arithmetic shown in the speed table.
FAQ
Got the hardware sorted? Now build on it.
You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.
Decide before you spend a thousand pounds
The AI Hardware course sizes your build properly — VRAM ladder, real bottlenecks, budget builds — and Pick the Right Model tells you what to run on it.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARLocal AI Hardware Requirements (2026): Complete Guide
- AI Hardware Requirements: CPU, GPU and RAM for Beginners
- AI RAM Requirements 2026: How Much for 7B, 13B, 70B Models?
- AI Server Build Under $1,500: Parts List and What Fits
- AMD GPU Not Supported by ROCm? HSA_OVERRIDE Values
- AMD MI50 32GB for Local LLMs: The Used VRAM King, Honestly
- AMD Ryzen AI Max+ 395 (Strix Halo) for Local AI 2026
- Apple M4 for Local AI: Mac Studio + MacBook Guide (2026)
- Benchmark Your Local AI Setup: tok/s, TTFT, VRAM
- Best GPU for AI Video Generation: By VRAM Tier (2026)
Comments (0)
No comments yet. Be the first to share your thoughts!