Best Coding LLM for 16GB VRAM (16GB VRAM)
The code models that fit a 16GB card now โ and the near-misses that do not
The best local coding LLM for 16GB VRAM in 2026 is Devstral Small 2 24B โ about 15GB at Q4_K_M, and Mistral reports 68.0% on SWE-bench Verified, the strongest agentic-coding score of anything that fits a 16GB card. The catch is headroom: ~15GB of weights leaves roughly a gigabyte for KV cache, so if your agent needs a long repo context run the MoE gpt-oss:20b or the dense Qwen 2.5 Coder 14B (~9GB) instead. The two models that would beat all three โ Qwen3.6 27B (~17GB) and Qwen3-Coder 30B (~19GB) โ miss 16GB by a hair, which is the entire argument for a 24GB card.
Models that fit in 16GB VRAM
Tested reference: RTX 4080 / RTX 4060 Ti 16GB. Figures are for Q4_K_M with a modest context window.
| Model | Size | Build | VRAM used | Speed |
|---|---|---|---|---|
| WINNERDevstral Small 2 24B Purpose-built agentic coder โ Mistral reports 68.0% SWE-bench Verified, Apache 2.0, 256K native context. At ~15GB it is a tight 16GB fit: set num_ctx deliberately and keep the card to itself. The tok/s figure is a reported RTX 3090 result at Q4 with a 64K context; no 16GB-card benchmark has been published. ollama pull devstral-small-2:24b | 24B (dense) | Q4_K_M | ~15GB | ~18โ22 tok/s (RTX 3090) |
| gpt-oss:20b OpenAIโs open-weight model, sized for 16GB cards. Its model card (arXiv 2508.10925) puts the 20B at 81.7% HumanEval, and MoE routing means it generates at small-model speed despite its size โ the fast, general-purpose half of this tier. ollama pull gpt-oss:20b | 20.9B (MoE, ~3.6B active) | MXFP4 | ~12โ16GB | ~40โ50 tok/s |
| Qwen 2.5 Coder 14B Still the sane default when context matters more than benchmark position โ ~7GB spare for KV cache means genuine repo-wide work, and no 2026 Qwen coder ships in the 14B class. ollama pull qwen2.5-coder:14b | 14B | Q4_K_M | ~9.0GB | ~34 tok/s |
| DeepSeek-Coder-V2 Lite 16B Only ~2.4B params fire per token, so it types faster than the dense 14B; DeepSeek publishes 90.2% HumanEval for it. ollama pull deepseek-coder-v2:16b | 16B (MoE, 2.4B active) | Q4_K_M | ~10GB | ~45 tok/s |
| Qwen3.8-27B (3-bit) Qwenโs newest 27B does fit 16GB โ but only at 3-bit, and only via the Unsloth GGUF repo; the default Ollama tag is the 18GB Q4 build. It is a vision-capable generalist, not a code specialist, so treat it as the experiment beside a real coder. unsloth/Qwen3.8-27B-GGUF ยท UD-Q3_K_XL | 27B (dense) | UD-Q3_K_XL GGUF | 12.5 GiB | no published figure |
| Qwen 2.5 Coder 7B The autocomplete half of a two-model setup โ small enough to sit resident beside the 14B and answer inline completions instantly. ollama pull qwen2.5-coder:7b-instruct-q5_K_M | 7B | Q5_K_M | ~5.5GB | ~48 tok/s |
What won't fit in 16GB
- โQwen3.6 27B (needs ~17GB) โ The dense 27B that scores 77.2% SWE-bench Verified on Qwenโs own model card misses a 16GB card by about a gigabyte at Q4_K_M โ the most frustrating near-miss in this tier, and the reason 24GB keeps winning the argument.
- โQwen3-Coder 30B A3B (needs ~19GB) โ MoE coder with a 256K native context; ~19GB at Q4_K_M on Ollama, so it is a 24GB pull.
- โQwen 2.5 Coder 32B (needs ~20GB) โ The strongest dense open coder (92.7% HumanEval) โ still a 24GB card away.
- โDevstral 2 123B (needs ~64โ80GB at 4-bit) โ The 72.2% SWE-bench Verified flagship. Open weights, but multi-GPU or workstation hardware.
How to fit more in 16GB
- โSet num_ctx explicitly before you point an agent at anything. Ollama defaults a modelโs context to roughly 2โ4K tokens, which Cline or Aider blows through in a few tool calls and then silently loops. 32K is the practical floor โ and on 16GB that means Devstral is the only thing on the card.
- โIf your work is repo-wide rather than single-file, take the smaller model on purpose: Qwen 2.5 Coder 14B at ~9GB leaves ~7GB of KV budget, and context you can actually feed beats parameters you cannot.
- โMoE is the 16GB cheat code โ gpt-oss:20b activates ~3.6B params per token and DeepSeek-Coder-V2 Lite ~2.4B, so both generate far faster than a dense 22โ24B squeezed in at Q4.
- โSplit the roles: a big model for chat and refactors, Qwen 2.5 Coder 7B or 3B for inline completion. Continue.dev and Cline both assign a model per role, and the small one is what makes the editor feel instant.
- โAfter loading, check that `ollama ps` says 100% GPU. Any CPU share means it spilled past 16GB and tok/s collapses โ on this tier the culprit is almost always context, not weights.
Quick start
curl -fsSL https://ollama.com/install.sh | shollama run devstral-small-2:24bollama run gpt-oss:20bGo from "it runs" to actually building
All 561 chapters โ running local models, RAG, agents, fine-tuning โ plus the Python Lab and every course added later.
Frequently asked questions
What is the best local coding model for 16GB VRAM?
Devstral Small 2 24B at Q4_K_M (~15GB) โ Mistral reports 68.0% on SWE-bench Verified, the best agentic-coding score of anything that fits. It is a tight fit, so if you need a long repo context run Qwen 2.5 Coder 14B (~9GB) or gpt-oss:20b instead. All three are a clear step up from the 7B coders an 8GB card is limited to.
Devstral Small 2 or gpt-oss:20b on a 16GB card?
Devstral if the job is agentic software engineering โ it was trained for the read-edit-test loop and carries the 68.0% SWE-bench Verified score Mistral reports. gpt-oss:20b if you want speed plus a general-purpose model: it is MoE (~3.6B active params per token) so it generates several times faster, and OpenAIโs model card puts the 20B at 81.7% HumanEval. Neither leaves room for the other, so pick one per session.
Is Qwen 2.5 Coder 14B still worth running in 2026?
Yes, and it is still the right pick for context-heavy work. It has been overtaken on agentic benchmarks, but at ~9GB it leaves ~7GB for KV cache on a 16GB card where Devstral leaves about one. Qwen has not shipped a 2026 coder in the 14B class โ Qwen3-Coder starts at 30B (~19GB) โ so the 14B remains the largest Qwen coder that fits comfortably.
Can 16GB VRAM run Qwen3.6 27B or the 32B coder?
Not at Q4. Qwen3.6 27B needs ~17GB and Qwen 2.5 Coder 32B ~20GB, so both spill into system RAM on a 16GB card and slow to a crawl. Qwen3.8-27B does fit at 3-bit (UD-Q3_K_XL, 12.5 GiB) if you want a 27B on this card, but you are trading precision for parameters. The full-quality 27โ32B class is what 24GB cards are for.
Related guides
Ready to Go Beyond Tutorials?
25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.
Was this helpful?