CodeLlama-13B: Technical Analysis
Code Llama 13B is the mid-size model in the family Meta released on 24 August 2023 with the paper Code Llama: Open Foundation Models for Code. It has 13 billion parameters, was trained on 16K-token sequences, and downloads as a 7.4 GB file with ollama run codellama:13b. It needs a 12 GB graphics card, not an 8 GB one. It is also three years old: newer code models score far higher and are smaller. This page covers what it needs, what the paper reported, and where it is still useful.
CodeLlama 13B hardware requirements
The default codellama:13b download is a 7.4 GB 4-bit file that needs about 11.8 GB of memory with a 4K context, so it fits a 12 GB graphics card with almost nothing to spare and does not fit an 8 GB card. Using the full 16K context raises that to about 21.8 GB, which is a 24 GB card. On CPU, 16 GB of system RAM runs the 4-bit file at 4K context.
Sizes are the download sizes on Ollama's codellama tag list. The memory columns are arithmetic, not measurements: file size, plus the key-value cache, plus a 1 GB allowance for compute buffers. Code Llama 13B inherits Llama 2's layout of 40 layers and a hidden size of 5,120 with no grouped-query attention, so the cache costs 2 × 40 × 5,120 × 2 bytes = 819 KB per token. That is 3.4 GB at 4,096 tokens and 13.4 GB at the full 16,384, which is why context length matters more here than the quantisation level.
| Quantisation | Ollama tag | Download | Memory at 4K | Memory at 16K | Fits |
|---|---|---|---|---|---|
| q2_K | codellama:13b-instruct-q2_K | 5.4 GB | ~9.8 GB | ~19.8 GB | 12 GB card at 4K |
| q3_K_M | codellama:13b-instruct-q3_K_M | 6.3 GB | ~10.7 GB | ~20.7 GB | 12 GB card at 4K |
| q4_0 (what codellama:13b pulls) | codellama:13b-instruct-q4_0 | 7.4 GB | ~11.8 GB | ~21.8 GB | 12 GB card at 4K, with almost nothing to spare; 24 GB card at 16K |
| Q4_K_M | codellama:13b-instruct-q4_K_M | 7.9 GB | ~12.3 GB | ~22.3 GB | 16 GB card at 4K; 24 GB card at 16K |
| Q5_K_M | codellama:13b-instruct-q5_K_M | 9.2 GB | ~13.6 GB | ~23.6 GB | 16 GB card at 4K; 24 GB card at 16K, just |
| Q6_K | codellama:13b-instruct-q6_K | 11 GB | ~15.4 GB | ~25.4 GB | 16 GB card at 4K, tight; 16K needs more than 24 GB |
| Q8_0 | codellama:13b-instruct-q8_0 | 14 GB | ~18.4 GB | ~28.4 GB | 24 GB card at 4K |
| fp16 | codellama:13b-instruct-fp16 | 26 GB | ~30.4 GB | ~40.4 GB | 32 GB card at 4K; 48 GB for 16K |
By graphics card
- 8 GB card: no 13B tag fits with a 4K context. Even the 5.4 GB q2_K file needs about 9.8 GB. Ollama will split the model between GPU and CPU, which works but is slow; CodeLlama 7B (3.8 GB) or a newer 7B code model is the better fit.
- 12 GB card: q4_0 at 4K context. Close other GPU applications first, or drop to q3_K_M for headroom.
- 16 GB card: Q4_K_M or Q5_K_M at 4K, or q4_0 with roughly 8K of context.
- 24 GB card: q4_0 through Q5_K_M with the full 16K context, or Q8_0 at 4K.
Ollama's context-length documentation gives a 4K default below 24 GiB of VRAM and 32K from 24 GiB, so on a 24 GB card it is worth setting the context to 16K explicitly, since that is this model's limit. The VRAM calculator runs the same sum for other models, and quantization explained covers what each level trades away.
The paper: Code Llama: Open Foundation Models for Code (arXiv 2308.12950)
The Code Llama paper is arXiv:2308.12950, by Baptiste Rozière, Jonas Gehring, Fabian Gloeckle and colleagues at Meta AI. It was submitted on 24 August 2023 and last revised on 31 January 2024; the current version also covers the 70B models. What it says, in the order most readers need it:
| Topic | What the paper and model card state |
|---|---|
| Basis | A family of code models built on Llama 2 |
| Variants | Code Llama (foundation), Code Llama - Python, Code Llama - Instruct |
| Sizes | 7B, 13B and 34B at launch; 70B added later and included in the revised paper |
| Training data | 500B tokens for Code Llama; a further 100B tokens of mostly Python for the Python variant |
| Context | Trained on sequences of 16K tokens; the abstract reports improvements on inputs of up to 100K tokens |
| Infilling | Supported by the 7B, 13B and 70B Code Llama and Code Llama - Instruct models; not by the Python variant or the 34B |
| Training period | January 2023 to July 2023 |
| 13B architecture | 40 layers, hidden size 5,120, 40 attention heads, vocabulary 32,016; 13,016,028,160 parameters in the published weights |
| Licence | Llama 2 Community License |
The official code is at github.com/meta-llama/codellama, which lists the unquantised 13B weights at 24 GB, and the Transformers-format weights are at codellama/CodeLlama-13b-hf.
Reported benchmarks
The paper reports HumanEval pass@1 of 36.0% for Code Llama 13B, 42.7% for the Instruct variant and 43.3% for the Python variant, with MBPP pass@1 of 47.0%, 49.4% and 49.0%. These are Meta's figures; this site has not run them.
| Model | HumanEval pass@1 | HumanEval pass@10 | MBPP pass@1 | MBPP pass@10 |
|---|---|---|---|---|
| Code Llama 7B | 33.5% | 59.6% | 41.4% | 66.7% |
| Code Llama 13B | 36.0% | 69.4% | 47.0% | 71.7% |
| Code Llama - Instruct 13B | 42.7% | 71.6% | 49.4% | 71.2% |
| Code Llama - Python 13B | 43.3% | 77.4% | 49.0% | 74.0% |
| Code Llama 34B | 48.8% | 76.8% | 55.0% | 76.2% |
Two things stand out. The step from 7B to 13B is small on HumanEval pass@1 (33.5% to 36.0%) and larger on MBPP (41.4% to 47.0%), while the step from 13B to 34B is much bigger on both. And within the 13B size, the Instruct and Python variants are about seven points ahead of the base model on HumanEval, so the base model is the one to pick only when you want raw completion or infilling.
The paper's multilingual table (MultiPL-E, pass@1) gives Code Llama 13B 39.1% on C++, 38.0% on Java, 34.2% on PHP, 29.6% on TypeScript, 27.3% on C# and 15.2% on Bash. Shell scripting is its weakest language by a wide margin.
Ollama setup and tags
Every tag below was checked against the Ollama tag list on 29 September 2026. One thing the short name hides: codellama:13b has the same ID as codellama:13b-instruct, so the default 13B download is the Instruct model, not the base model.
# Install Ollama (Linux; macOS and Windows use the installer from ollama.com)
curl -fsSL https://ollama.com/install.sh | sh
# Instruct, 4-bit, 7.4 GB — chat-style coding questions
ollama run codellama:13b
# Base model, 7.4 GB — completion and fill-in-the-middle
ollama pull codellama:13b-code
# Python specialist, 7.4 GB
ollama pull codellama:13b-python
# Smaller file for headroom on a 12 GB card, 6.3 GB
ollama pull codellama:13b-instruct-q3_K_MFill-in-the-middle
Infilling is the feature that still gives Code Llama a job in editor autocomplete. The base model takes a prefix and a suffix and writes what goes between them. Ollama's library page documents the prompt format with the 7B tag; the same format applies to 13b-code:
ollama run codellama:13b-code '<PRE> def compute_gcd(x, y): <SUF>return result <MID>'For wiring a local model into an editor, see the Continue.dev with Ollama guide, and for current autocomplete picks the local autocomplete models comparison.
Base, Instruct or Python: which 13B to pull
| Variant | Ollama tag | Use it for | Infilling |
|---|---|---|---|
| Code Llama - Instruct | codellama:13b or codellama:13b-instruct | Asking for code or explanations in plain language | Yes |
| Code Llama (base) | codellama:13b-code | Continuing code from a prompt; editor autocomplete | Yes |
| Code Llama - Python | codellama:13b-python | Python completion; it is not tuned to follow instructions | No |
The official repository is explicit that the base and Python models “are not fine-tuned to follow instructions” and should be prompted so that the answer is the natural continuation of the prompt. If a Code Llama model answers a question with more code comments instead of an answer, that is usually the reason.
What replaced CodeLlama 13B
For new work, use a newer code model. Qwen2.5-Coder 7B is a smaller download than CodeLlama 13B (4.7 GB against 7.4 GB), is Apache 2.0 licensed, has a 131,072-token context, and scores more than twice as high on HumanEval in its own technical report.
The scores below are from the instruct-model table of the Qwen2.5-Coder technical report (arXiv:2409.12186), published by Alibaba's Qwen team, who evaluated CodeLlama alongside their own models. Their CodeLlama-13B-Instruct HumanEval figure (40.2) differs slightly from the 42.7% in Meta's paper because the evaluation setups differ. Download sizes are from Ollama.
| Model | HumanEval | HumanEval+ | MBPP | MBPP+ | Ollama tag and size |
|---|---|---|---|---|---|
| CodeLlama-13B-Instruct | 40.2 | 32.3 | 60.3 | 51.1 | codellama:13b, 7.4 GB |
| Qwen2.5-Coder-7B-Instruct | 88.4 | 84.1 | 83.5 | 71.7 | qwen2.5-coder:7b, 4.7 GB |
| Qwen2.5-Coder-14B-Instruct | 89.6 | 87.2 | 86.2 | 72.8 | qwen2.5-coder:14b, 9.0 GB |
The site has pages for Qwen 2.5 Coder 7B, Qwen 2.5 Coder 14B, Qwen3-Coder-Next and Devstral; the current shortlist is on best local coding models, and the Ollama coding-model guide sorts them by VRAM tier.
Reasons to keep CodeLlama 13B are narrow. You may have a fine-tune or an evaluation built on it and need the same weights; you may be reproducing the paper; or a tool you use is configured for Code Llama's infill tokens. Within the family, the 34B model is the better use of a 24 GB card, though it does not support infilling.
Frequently asked questions
Will CodeLlama 13B run on an 8 GB graphics card?
Not entirely on the card. The 4-bit file alone is 7.4 GB, and a 4K context adds a 3.4 GB cache. Ollama will place what fits on the GPU and run the rest on the CPU, which is slower than a model that fits completely. A 7B model is the right size for 8 GB.
Is the context window 16K or 100K?
The model was trained on 16K-token sequences and its configuration sets a 16,384-token limit, which is also what Ollama lists for every 13B tag. The 100K figure comes from the paper's finding that the models keep improving on inputs up to that length. On local hardware the cache is the constraint: 16K already costs 13.4 GB on top of the model.
Can I use CodeLlama 13B commercially?
The model card says Code Llama is “intended for commercial and research use”, and the weights ship under the Llama 2 Community License. That is Meta's own licence with an acceptable use policy, not an open-source licence such as Apache 2.0 or MIT, so read the licence text before shipping a product on it.
How does a 13B model compare with a 70B one for coding?
Within Code Llama, the paper reports HumanEval pass@1 of 36.0% for the 13B base model and 53.0% for the 70B base model, and 42.7% against 67.8% for the Instruct variants. The 70B is a 39 GB download on Ollama against 7.4 GB. Model generation now matters more than size, though: the 7B Qwen2.5-Coder in the table above outscores Code Llama 70B Instruct in Alibaba's evaluation. The guide to model size for coding goes through the tiers.
Is 13B worth it over CodeLlama 7B?
By the paper's numbers the gain is 2.5 points of HumanEval pass@1 and 5.6 points of MBPP pass@1 for the base models, for nearly twice the download (7.4 GB against 3.8 GB) and a cache that is about twice as large per token. On a 12 GB card the 13B fits; on anything smaller the 7B is the practical choice.
Sources
- Code Llama: Open Foundation Models for Code (arXiv:2308.12950)Training data, context length, infilling, and every Code Llama score on this page
- CodeLlama-13b-hf model cardCapabilities, training period, intended use and licence
- Official Code Llama repositoryReference inference code, weight sizes and prompting notes
- Ollama codellama tag listEvery tag, its ID and its download size
- Ollama codellama library pageThe infill prompt format
- Qwen2.5-Coder technical report (arXiv:2409.12186)The comparison between CodeLlama and Qwen2.5-Coder instruct models
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Related Guides
Continue your local AI journey with these comprehensive guides
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Written by the Local AI Master Team
The team behind Local AI Master
We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.