Codestral 22B: Code Generation Analysis
Updated: September 29, 2026
Codestral 22B (Codestral-22B-v0.1) is the open-weight code model Mistral AI released on 29 May 2024. It has 22.2 billion parameters and a 32K context window, was trained on more than 80 programming languages, and works in two modes: instruct, and fill-in-the-middle for editor autocomplete. The default Ollama build is a 13 GB download that runs on a 16 GB graphics card. Its weights are under the Mistral AI Non-Production Licence, which allows research, testing and personal use but not use in the course of paid work. Everything below is taken from Mistral's announcement, the model card, the licence text and the published file sizes; this site has not benchmarked the model.
Codestral 22B hardware requirements
Codestral 22B at Q4_K_M is a 13.3 GB file. It runs on a 16 GB card with a short context and comfortably on a 24 GB card, where there is room for the full 32K window. A 12 GB card can hold the 8.3 GB Q2_K build. Q8_0 is 23.6 GB and needs a 32 GB card. The arithmetic is parameters × bits per weight ÷ 8: 22.25B × 4.8 bits ÷ 8 = 13.3 GB at Q4_K_M.
File sizes are from bartowski/Codestral-22B-v0.1-GGUF and the tag sizes from Ollama's tag list. The context cache is arithmetic from the model's params.json: 56 layers × 8 key-value heads × 128 dimensions, stored twice (keys and values) at 2 bytes, is 229 KB per token — 0.9 GB at 4K context, 1.9 GB at 8K and 7.5 GB at the full 32K. “Memory needed” is file + the 8K cache + a 1 GB allowance for buffers.
| Quantisation | GGUF file | Ollama tag (listed size) | Memory needed (8K context) | Fits |
|---|---|---|---|---|
| Q2_K | 8.27 GB | 22b-v0.1-q2_K (8.3GB) | ~11.2 GB | 12 GB card (RTX 3060 12GB, RTX 4070) |
| Q3_K_M | 10.76 GB | 22b-v0.1-q3_K_M (11GB) | ~13.7 GB | 16 GB card |
| Q4_K_M | 13.34 GB | 22b-v0.1-q4_K_M (13GB) | ~16.2 GB | 16 GB card only at 4K context; 24 GB comfortably |
| Q5_K_M | 15.72 GB | 22b-v0.1-q5_K_M (16GB) | ~18.6 GB | 24 GB card (RTX 3090, RTX 4090) |
| Q6_K | 18.25 GB | 22b-v0.1-q6_K (18GB) | ~21.2 GB | 24 GB card at 8K context |
| Q8_0 | 23.64 GB | 22b-v0.1-q8_0 (24GB) | ~26.5 GB | Does not fit 24 GB; 32 GB card (RTX 5090) |
With the full 32K context, Q4_K_M needs about 13.3 + 7.5 + 1 = 21.8 GB, which is why 24 GB is the comfortable tier. On a Mac the same totals apply to unified memory, and without a GPU the file loads into system RAM, so 32 GB of RAM covers Q4_K_M. Speed depends on the hardware and is not quoted here because this site has not measured it. How the build names trade size for quality is covered in quantization explained, and the Ollama RAM and VRAM table lists other models by size.
Licence: what you may use it for
Codestral 22B is not open source. The model card on Hugging Face carries the licence tag mnpl and links to the Mistral AI Non-Production Licence, version 0.1. The repository was last updated in July 2025 and still carries that licence.
| Use | Under the MNPL |
|---|---|
| Research, testing and evaluation | Allowed, in a non-production environment |
| Personal projects with no commercial connection | Allowed |
| Coding assistant for your day job | Not allowed — the licence excludes “any usage by individuals employed in companies in the context of their daily tasks” |
| Hosting it as a service, paid or free | Not allowed |
| Fine-tuning | Allowed, and the fine-tune stays under the same restrictions |
Mistral's announcement says the licence “means that you can use it for research and testing purposes” and that “commercial licenses are also available on demand by reaching out to the team”. The licence text also states that the model's outputs are not treated as derivatives of the model. This is a summary, not legal advice; read the licence before relying on it.
An earlier version of this page said Codestral had been relicensed under Apache 2.0. That could not be confirmed against Mistral's model card or licence file and has been removed. If you need a permissively licensed code model, see the alternatives below.
Model specifications
| Spec | Codestral-22B-v0.1 |
|---|---|
| Developer | Mistral AI |
| Released | 29 May 2024 |
| Parameters | 22,247,282,688 (BF16 weights) |
| Context length | 32K (32,768 positions) |
| Architecture | Dense transformer: 56 layers, width 6144, 48 attention heads, 8 key-value heads |
| Vocabulary | 32,768 tokens |
| Languages | 80+ programming languages; Mistral names Python, Java, C, C++, JavaScript, Bash, Swift and Fortran |
| Modes | Instruct and fill-in-the-middle (FIM) |
| Moderation | None — the card says the model “does not have any moderation mechanisms” |
| Licence | Mistral AI Non-Production Licence 0.1 |
Fill-in-the-middle is what sets a completion model apart from a chat model. The editor sends the code before the cursor and the code after it, and the model writes what belongs in between. That is the mode an autocomplete plug-in uses; instruct mode is for asking questions about code or requesting a function.
Installation
The quickest route is Ollama. codestral, codestral:22b and codestral:v0.1 are the same 13 GB build. All 17 tags on ollama.com/library/codestral/tags list a 32K context.
# Default build, 13 GB
ollama pull codestral:22b
ollama run codestral:22b
# Smaller build for a 12 GB card, 8.3 GB
ollama pull codestral:22b-v0.1-q2_K
# Higher-precision build for a 24 GB card, 18 GB
ollama pull codestral:22b-v0.1-q6_KOriginal weights with mistral-inference
Mistral recommends its own mistral-inference package for the unquantised weights. These two commands are from the model card:
pip install mistral_inference
mistral-chat $HOME/mistral_models/Codestral-22B-v0.1 --instruct --max_tokens 256The BF16 weights are about 44 GB (22.25B parameters × 2 bytes), so that route needs a 48 GB card or two 24 GB cards. For editor autocomplete, the Continue.dev with Ollama guide covers pointing a VS Code or JetBrains plug-in at a local model, and best local autocomplete models compares the options.
What Mistral published about performance
Mistral's announcement evaluates Codestral on HumanEval and MBPP for Python generation, CruxEval for predicting program output, RepoBench for long-range repository-level completion, Spider for SQL, and HumanEval in six further languages. Its fill-in-the-middle results are compared against DeepSeek Coder 33B in Python, JavaScript and Java. The scores are published as chart images on that page, so they are not reproduced here; read them at the source.
The claim Mistral makes in words is about context: “With its larger context window of 32k (compared to 4k, 8k or 16k for competitors), Codestral outperforms all other models in RepoBench”. That was a statement about the models available in May 2024.
The one independent figure on the page is from a JetBrains researcher, who reported a Kotlin-HumanEval pass rate of 73.75 for Codestral at temperature 0.2, against 72.05 for GPT-4-Turbo and 54.66 for GPT-3.5-Turbo.
What has replaced it
Codestral 22B is the first Codestral and the only full-size one with downloadable weights. Mistral has since announced Codestral 25.01 (13 January 2025), which it describes as “about 2 times faster” than the original, and Codestral 25.08 (30 July 2025). Neither appears among Mistral's repositories on Hugging Face, where the only Codestral entries are this model and the 7B Mamba-Codestral, so they are hosted models rather than something you can pull.
For local use in 2026 the practical replacements are newer open-weight code models with permissive licences:
- Devstral — Mistral's own later code model. Devstral Small is Apache-2.0, 23.6B parameters, and the Ollama
devstral:24btag is 14 GB with a 128K context. - Qwen 2.5 Coder 32B — Apache-2.0, for a 24 GB card.
- Qwen3-Coder-Next and the best local AI coding models list for the current picks.
Codestral 22B is still a reasonable choice for personal autocomplete on a 16 GB card, where its fill-in-the-middle mode is the point. For anything connected to paid work, the licence rules it out before quality comes into it.
Frequently asked questions
What is the context length of Codestral 22B?
32K tokens. Mistral's announcement says 32k, the configuration file sets 32,768 positions, and every Ollama tag lists a 32K context window. Filling it costs about 7.5 GB of memory on top of the model file.
How many parameters does Codestral have?
22,247,282,688, which Mistral rounds to 22B. It is a dense model, so all of them are used for every token.
Can I use Codestral 22B at work?
Not under the licence it ships with. The Non-Production Licence limits use to testing, research, personal and evaluation purposes, and defines personal use so that it excludes work done as an employee. Mistral offers separate commercial licences on request.
Will it run on a 12 GB or 16 GB card?
On 12 GB, only the 8.3 GB Q2_K build fits with a useful context, and 2-bit quantisation costs quality. On 16 GB, Q3_K_M fits with room to spare and Q4_K_M fits if you keep the context to about 4K. If you are choosing a model for that tier rather than this model specifically, see the Ollama coding-model guide.
Is Codestral the same as Devstral?
No. Codestral is a completion and instruct model from May 2024 under a non-production licence. Devstral is a later Mistral model aimed at agentic coding, and Devstral Small is Apache-2.0.
Sources
- Mistral AI: Codestral announcement — release date, context, languages, licence statement
- Codestral-22B-v0.1 model card — parameter count, usage modes, limitations
- Mistral AI Non-Production Licence 0.1 — the licence text
- Ollama tag list — every tag with its download size
- bartowski GGUF files — the file sizes in the hardware table
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Written by the Local AI Master Team
The team behind Local AI Master
We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.
Related Guides
Continue your local AI journey with these comprehensive guides
- PILLARLocal AI Models Directory: Every Model Compared
- Alpaca 7B: Stanford\
- Amazon Chronos: Time Series Forecasting Models (Complete Guide)
- Anima 2.9B on 8GB: The Anime Model Taking Over Civitai
- Aquila 7B by BAAI: Chinese-English Bilingual (FlagAI)
- Baichuan2-13B: Chinese LLM | 59% CMMLU, Bilingual, Free License 2026
- Bark by Suno AI: Open-Source Text-to-Audio Generation Guide
- ChatGLM3-6B: Tsinghua Chinese AI | Code Interpreter, 6GB RAM 2026
- Claude 3 Opus Review: Benchmarks, Pricing & API Guide 2026
- Claude 3 Sonnet Review: Benchmarks, API Pricing & Alternatives 2026
Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide
No spam. Unsubscribe with one click.
Found your model? Now build something with it.
25 hands-on courses — RAG, agents, fine-tuning — all running locally. First chapter free, no card.