DeepSeek Coder V2 236B
Hardware Requirements, Cost and Local Setup
Updated: September 22, 2026
DeepSeek Coder V2 is DeepSeek's June 2024 open-weight coding model. It ships in two sizes: the full model with 236B total parameters, 21B active per token, and a Lite model with 16B total and 2.4B active. Both are Mixture-of-Experts, both carry a 128K context window, and both are trained on 338 programming languages. The full model does not fit any single consumer GPU; the Lite model does. Everything on this page is taken from the model card, the GitHub README, the paper and the published GGUF and Ollama file sizes, which are linked in each section.
DeepSeek Coder V2 hardware requirements
DeepSeek-Coder-V2-Lite (16B total, 2.4B active) runs on a single 24 GB card at Q8_0 with its full 128K context, and on a 12 GB card at Q4_K_M with a short context. The full DeepSeek-Coder-V2 (236B total, 21B active) does not fit any single consumer GPU: the smallest published GGUF is 52.7 GB, the standard Q4_K_M is 142.5 GB, and DeepSeek's own model card says BF16 inference needs “80GB*8 GPUs”. Both models are Mixture-of-Experts, so the active-parameter count sets speed, not memory — every expert has to be resident in VRAM or system RAM before the first token. The MoE VRAM calculator does the same arithmetic for any expert count.
File sizes are the shard sums reported by the Hugging Face API for bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF and bartowski/DeepSeek-Coder-V2-Instruct-GGUF; the fp16 and larger 236B rows come from the sizes on Ollama's tag list. KV cache is arithmetic from each config.json: Coder V2 uses multi-head latent attention, so llama.cpp caches (512 + 64) values per token per layer at 2 bytes — 27 layers for Lite = 31 KB per token (0.5 GB at 16K context, 4.1 GB at 128K), 60 layers for the 236B = 69 KB per token (1.1 GB at 16K, 9.1 GB at 128K). “Memory needed” below is file + 16K KV cache + a 1 GB compute-buffer allowance.
Lite 16B: what fits a single card
| Quantization | GGUF file | Memory needed (16K ctx) | Fits |
|---|---|---|---|
| IQ4_XS | 8.57 GB | ~10.1 GB | 12 GB card (RTX 3060 12GB, RTX 5070) |
Q4_K_M (Ollama 16b is the 8.9 GB q4_0) | 10.4 GB | ~11.9 GB | 12 GB card only at short context; 16 GB comfortably |
| Q5_K_M | 11.9 GB | ~13.4 GB | 16 GB card (RTX 4060 Ti 16GB, 5060 Ti 16GB, 4080) |
| Q6_K | 14.1 GB | ~15.6 GB | 16 GB card at 16K; 24 GB for longer context |
| Q8_0 | 16.7 GB | ~18.2 GB (~21.8 GB at the full 128K) | 24 GB card (RTX 3090/4090) with the whole 128K context |
| fp16 | 31 GB (Ollama tag) | ~32.5 GB | Does not fit 24 GB or a 32 GB RTX 5090 with context; 48 GB |
236B: 48–96 GB of VRAM, or system RAM offload
| Quantization | GGUF file | Fully in VRAM | With a 24 GB card + system RAM (experts offloaded) |
|---|---|---|---|
| IQ1_M | 52.7 GB | 64 GB: 2× RTX 5090, or 3× 24 GB cards | 64 GB RAM |
| IQ2_XS | 68.7 GB | 96 GB: RTX PRO 6000 Blackwell, or 4× 24 GB | 96 GB RAM (64 GB plus the card's 24 GB works, with little headroom) |
Q2_K (Ollama 236b q2_K is 86GB) | 85.9 GB | 96 GB, with under 10 GB left for context | 96 GB RAM |
| Q3_K_M | 112.7 GB | 128 GB: 2× RTX PRO 6000, or 4× 32 GB | 128 GB RAM |
Q4_K_M (Ollama 236b = 133GB q4_0) | 142.5 GB | 160 GB+: 2× H100 80GB, or 2× RTX PRO 6000 (192 GB) | 192 GB RAM (128 GB plus the card's 24 GB is 152 GB total, about 7 GB of headroom) |
| Q5_K_M / Q6_K / Q8_0 (Ollama) | 167 / 194 / 251 GB | 3–4× H100 80GB | 192 / 256 / 256 GB RAM |
| BF16 (official weights) | 472 GB (Ollama fp16 tag) | “80GB*8 GPUs are required” — DeepSeek model card | Not a workstation configuration |
The RAM column is file size + 1 GB buffer + the 1.1 GB 16K KV cache, rounded up to the next common DDR5 kit with room for the OS; offloading keeps the attention layers and cache on the card and streams the routed experts from RAM over PCIe, so it works but is bandwidth-bound. Ollama's 236b tag is q4_0 (133 GB), not Q4_K_M. For a coding model that stays on one 24 GB card at full precision, compare Qwen3-Coder-Next and the best models for a 24 GB GPU; for the newer DeepSeek generation see DeepSeek V4 and its hardware requirements by memory tier. How the K-quant names trade size for quality is covered in quantization explained.
Model specifications
Figures below are from the DeepSeek-Coder-V2-Instruct model card and the paper DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence (DeepSeek-AI, submitted 17 June 2024). The model is a Mixture-of-Experts continuation of a DeepSeek-V2 checkpoint, further trained on 6 trillion tokens; the paper reports the language count growing from 86 to 338 and the context from 16K to 128K.
| Spec | DeepSeek-Coder-V2 (236B) | DeepSeek-Coder-V2-Lite (16B) |
|---|---|---|
| Total parameters | 236B | 16B |
| Active parameters per token | 21B | 2.4B |
| Context length | 128K | 128K |
| Programming languages | 338 | 338 |
| Variants on Hugging Face | Base, Instruct | Lite-Base, Lite-Instruct |
| Official weights | BF16 (“80GB*8 GPUs are required”) | BF16 (31 GB as an fp16 GGUF) |
| Licence | Code under MIT; weights under the DeepSeek Model Agreement, which the README says “supports commercial use” for Base and Instruct | |
Reported benchmarks
These are DeepSeek's own numbers, copied from the table in the DeepSeek-Coder-V2 README for the two Instruct models. This site has not run these benchmarks; treat them as vendor-reported. The SWE-Bench column is the one to look at if you want an agentic coder rather than a completion model — the Lite model scores 0.0 there.
| Benchmark | DeepSeek-Coder-V2-Instruct (236B) | DeepSeek-Coder-V2-Lite-Instruct (16B) |
|---|---|---|
| HumanEval | 90.2 | 81.1 |
| MBPP+ | 76.2 | 68.8 |
| LiveCodeBench | 43.4 | 24.3 |
| Aider | 73.7 | 44.4 |
| SWE-Bench | 12.7 | 0.0 |
Installation
The quickest route is Ollama. deepseek-coder-v2 with no tag is the Lite Instruct model at q4_0 (8.9 GB), which is what most people should start with. The tags below are the ones that matter; the full list with every quantization is on ollama.com/library/deepseek-coder-v2/tags.
| Ollama tag | Download | Notes |
|---|---|---|
deepseek-coder-v2 (latest / lite / 16b) | 8.9 GB | Lite Instruct, q4_0. Same size and 160K context listing as 16b-lite-instruct-q4_0. |
deepseek-coder-v2:16b-lite-instruct-q8_0 | 17 GB | Lite at Q8_0; the 24 GB-card pick from the hardware table. |
deepseek-coder-v2:236b | 133 GB | Full model, q4_0 (not Q4_K_M). |
deepseek-coder-v2:236b-instruct-q2_K | 86 GB | Smallest 236B tag Ollama publishes; 96 GB of VRAM or RAM. |
Lite 16B with Ollama (one GPU)
# Lite Instruct at q4_0 (8.9 GB) — fits a 12 GB card at short context
ollama pull deepseek-coder-v2:16b
ollama run deepseek-coder-v2:16b
# Lite Instruct at Q8_0 (17 GB) — the 24 GB-card pick, room for the full 128K context
ollama pull deepseek-coder-v2:16b-lite-instruct-q8_0Full 236B with Ollama (multi-GPU or RAM offload)
# q4_0, 133 GB on disk — see the 236B table above for what memory this needs
ollama pull deepseek-coder-v2:236b
# q2_K, 86 GB — the smallest 236B tag Ollama publishes
ollama pull deepseek-coder-v2:236b-instruct-q2_Kllama.cpp with a specific GGUF
If you want a quantization Ollama does not carry (the IQ1_M and IQ2_XS 236B files, or IQ4_XS for Lite), llama.cpp's quick start loads a Hugging Face repo directly with -hf; the GGUF repos are the two bartowski links in the hardware section.
# Lite Instruct GGUFs (pick the quant file on the repo page)
llama serve -hf bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF
# 236B Instruct GGUFs — the IQ1_M (52.7 GB) and IQ2_XS (68.7 GB) files live here
llama serve -hf bartowski/DeepSeek-Coder-V2-Instruct-GGUFThe -hf syntax is from the llama.cpp README quick start. Which coding model to pull for each VRAM tier is covered in the Ollama coding-model guide.
Frequently asked questions
Which size should I run, the 236B or the 16B Lite?
Run the Lite model unless you already own 48–96 GB of VRAM or 96 GB+ of system RAM. Lite at Q8_0 is a 16.7 GB file that fits a 24 GB card with the whole 128K context; the smallest 236B file anyone publishes is 52.7 GB and Ollama's smallest 236B tag is 86 GB. DeepSeek's own table puts Lite at 81.1 HumanEval against 90.2 for the full model, so you give up some quality for a model that actually fits. The DeepSeek Coder V2 16B page covers the Lite model on its own.
Can I use it commercially? What is the licence?
The repository code is MIT-licensed (LICENSE-CODE) and the weights are under the DeepSeek Model Agreement (LICENSE-MODEL). The README states that the series, Base and Instruct, “supports commercial use”. Read the model agreement yourself before shipping anything; it is a custom licence, not Apache-2.0 or MIT.
How long is the context window, and what does it cost in memory?
128K tokens for both sizes, per the model card. Because Coder V2 uses multi-head latent attention the KV cache is small for a model this size: about 31 KB per token for Lite and 69 KB per token for the 236B, which works out to 4.1 GB and 9.1 GB respectively at the full 128K. That is why Lite at Q8_0 (16.7 GB) still fits a 24 GB card with the whole window. Ollama's tag list shows its own per-tag context defaults, so set the context length explicitly if you need the full 128K.
How does it compare with newer local coding models?
DeepSeek Coder V2 dates from June 2024 and it shows in the agentic benchmarks: DeepSeek reports SWE-Bench 12.7 for the 236B and 0.0 for Lite. Qwen2.5-Coder-32B-Instruct, released a few months later (technical report September 2024), is a 32.5B dense model under Apache-2.0 with a 131,072-token context, and its Ollama 32b tag is a 20 GB download — so it sits on a single 24 GB card while the 236B needs several. The site's current picks are on the best local coding models page, with individual pages for Qwen 2.5 Coder 32B, Qwen3-Coder-Next and Devstral. Coder V2 is still worth running if you want an MIT-code, commercially usable MoE model that is cheap per token, but it is no longer the first thing to try.
Base or Instruct, and which Ollama tag is which?
Instruct is the chat-tuned model you want for an assistant or an editor plugin; Base is the raw completion model for fill-in-the-middle or fine-tuning. On Ollama, the untagged deepseek-coder-v2 (and 16b, lite) is the 8.9 GB Lite Instruct q4_0; the Base variants are explicit tags such as 16b-lite-base-q4_0 and 236b-base-q4_0. Hugging Face carries all four: DeepSeek-Coder-V2-Base, -Instruct, -Lite-Base and -Lite-Instruct.
Sources and further reading
Official documentation
- •DeepSeek-Coder-V2 on GitHubREADME with the benchmark table, licence section and inference notes
- •DeepSeek-Coder-V2-Instruct model cardParameter counts, context length, licence and the “80GB*8 GPUs” note
- •DeepSeek-Coder-V2-Lite-Instruct model cardThe 16B model, with the config.json the KV-cache arithmetic comes from
- •DeepSeek-Coder-V2 paper (arXiv:2406.11931)Training data, architecture and evaluation methodology
Running it
- •Ollama tag listEvery published quantization with its download size
- •bartowski GGUF files (236B)The IQ1_M through Q4_K_M files sized in the hardware table
- •llama.cppCPU and mixed CPU/GPU inference; the expert-offload path for the 236B
- •vLLMMulti-GPU serving for the BF16 weights
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Written by the Local AI Master Team
The team behind Local AI Master
We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.
Related Guides
Continue your local AI journey with these comprehensive guides
- PILLARLocal AI Models Directory: Every Model Compared
- Alpaca 7B: Stanford\
- Amazon Chronos: Time Series Forecasting Models (Complete Guide)
- Anima 2.9B on 8GB: The Anime Model Taking Over Civitai
- Aquila 7B by BAAI: Chinese-English Bilingual (FlagAI)
- Baichuan2-13B: Chinese LLM | 59% CMMLU, Bilingual, Free License 2026
- Bark by Suno AI: Open-Source Text-to-Audio Generation Guide
- ChatGLM3-6B: Tsinghua Chinese AI | Code Interpreter, 6GB RAM 2026
- Claude 3 Opus Review: Benchmarks, Pricing & API Guide 2026
- Claude 3 Sonnet Review: Benchmarks, API Pricing & Alternatives 2026
Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide
No spam. Unsubscribe with one click.
Found your model? Now build something with it.
25 hands-on courses — RAG, agents, fine-tuning — all running locally. First chapter free, no card.