Phi-3 Small 7B: sizes, requirements and why it is not on Ollama
Updated: September 29, 2026
Phi-3 Small is Microsoft's 7.4 billion parameter model from May 2024, and it is the one Phi-3 size that Ollama does not carry. There is no phi3:small tag. The weights are 14.8 GB at bf16 and the int4 ONNX build is 5.0 GB, so it runs on a 24 GB card with Transformers or on an 8 GB card with ONNX Runtime.
If you came here for Ollama sizes: phi3:mini is a 2.2 GB download and runs on a machine with 8 GB of RAM, and phi3:medium is 7.9 GB. Both have been replaced by Phi-4 and Phi-4 Mini, which are on Ollama.
Phi-3 Small hardware requirements
A 24 GB NVIDIA card for the original weights, or an 8 GB card for the int4 ONNX build. The arithmetic is parameters × bits per weight: 7.39 billion × 16 bits ÷ 8 is 14.8 GB, which is the size of the safetensors files on Hugging Face. The 5.0 GB int4 build works out to 5.4 bits per weight.
| Build | Size | Runtime | What fits it |
|---|---|---|---|
| Original weights (bf16 safetensors) | 14.8 GB | Transformers with trust_remote_code | A 24 GB card. A 16 GB card holds the weights with about 1 GB left for context. |
| ONNX fp16 (CUDA) | 15.6 GB | ONNX Runtime | A 24 GB card. |
| ONNX int4 (CUDA) | 5.0 GB | ONNX Runtime | An 8 GB card. |
There is a catch with the original weights. The model card says Phi-3 Small uses flash attention 2 and Triton blocksparse attention by default, “which requires certain types of GPU hardware to run”, and the only GPU types Microsoft lists are the NVIDIA A100, A6000 and H100. Microsoft's sample code asserts that CUDA is available. For anything else, including CPUs and Windows machines with AMD or Intel graphics, the card points to the ONNX builds.
Sizes are the folder totals reported by the Hugging Face API for microsoft/Phi-3-small-8k-instruct-onnx-cuda. The “what fits” column is arithmetic, not a measurement. To size other models, use the VRAM calculator.
Phi-3 model sizes on Ollama, and RAM for phi3:mini
phi3:mini is 2.2 GB and needs about 5 GB of free memory at a 4K context, so it runs on a machine with 8 GB of RAM or a 6 GB graphics card. phi3:medium is 7.9 GB. Ollama's library describes Phi-3 as Mini and Medium only; Small is absent.
| Model | Ollama tag | Download | Note |
|---|---|---|---|
| Phi-3 Mini (3.8B) | phi3:mini | 2.2 GB | Same file as phi3 and phi3:3.8b. It is the q4_0 of the 128K variant. |
| Phi-3 Mini 4K | phi3:mini-4k | 2.4 GB | Q4_K_M of the 4K variant. |
| Phi-3 Mini, Q8_0 | phi3:3.8b-mini-4k-instruct-q8_0 | 4.1 GB | |
| Phi-3 Mini, fp16 | phi3:3.8b-mini-4k-instruct-fp16 | 7.6 GB | |
| Phi-3 Small (7B) | none | n/a | Ollama does not publish it. |
| Phi-3 Medium (14B) | phi3:medium | 7.9 GB | Same file as phi3:14b. |
| Phi-3 Medium, Q8_0 | phi3:14b-medium-4k-instruct-q8_0 | 15 GB | |
| Phi-3 Medium, fp16 | phi3:14b-medium-4k-instruct-fp16 | 28 GB |
The 5 GB figure for phi3:mini is the 2.2 GB file, plus 1.6 GB of KV cache for 4,096 tokens, plus a 1 GB allowance for buffers. Phi-3 Mini does not use grouped-query attention, so its cache is large for a model this size: 32 layers × 3072 × 2 × 2 bytes is 393 KB per token. That matters because Ollama lists phi3:mini with a 128K context. Filling all of it would take 51.5 GB of cache, so set the context length to what you need.
All sizes are from Ollama's phi3 tag list. The individual pages for Phi-3 Mini 3.8B and Phi-3 Medium 14B go further, and the Ollama system requirements guide covers memory by model size.
Specifications
From the model card and the config.json. The family is described in the Phi-3 Technical Report (arXiv:2404.14219, April 2024), which covers the 3.8B, 7B and 14B models together.
| Field | Phi-3 Small |
|---|---|
| Parameters | 7,392,274,432 (7.4B) |
| Architecture | Dense decoder-only transformer with alternating dense and blocksparse attention |
| Layers | 32 |
| Hidden size | 4096 |
| Attention heads | 32, with 8 key/value heads |
| Vocabulary | 100,352 (tiktoken) |
| Context length | 8K, with a separate 128K variant |
| Training data | 4.8 trillion tokens, cutoff October 2023 |
| Training run | 1024 H100-80G GPUs for 18 days |
| Weights released | 21 May 2024 |
| Licence | MIT |
Two of these set Small apart from its siblings. It uses the tiktoken tokenizer with a 100,352-token vocabulary, where Mini and Medium use a 32,064-token vocabulary. And it alternates dense and blocksparse attention layers, where the other two are standard transformers. Its config declares its own model type, phi3small. That is the reason the usual GGUF route did not pick it up: a llama.cpp request for conversion support, issue 8241, was closed as not planned in July 2024.
How to run Phi-3 Small
With Transformers on a supported NVIDIA GPU. The model card asks for tiktoken and triton and for trust_remote_code=True, because the model code ships in the repository. This is Microsoft's sample, shortened:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
model_id = "microsoft/Phi-3-small-8k-instruct"
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
trust_remote_code=True,
)
assert torch.cuda.is_available(), "This model needs a GPU to run ..."
device = torch.cuda.current_device()
model = model.to(device)
tokenizer = AutoTokenizer.from_pretrained(model_id)
pipe = pipeline("text-generation", model=model, tokenizer=tokenizer, device=device)
messages = [{"role": "user", "content": "What about solving an 2x + 3 = 7 equation?"}]
output = pipe(messages, max_new_tokens=500, return_full_text=False, do_sample=False)
print(output[0]["generated_text"])On other hardware, use the ONNX builds with ONNX Runtime GenAI. Microsoft publishes int4 builds for DirectML, CUDA, and CPU and mobile, plus an fp16 CUDA build. If what you want is a small Microsoft model behind ollama run, that is not this model; pull one of the models in the table below.
Benchmarks Microsoft reported
Copied from the Phi-3-small-8k-instruct model card. Microsoft ran every column through the same internal pipeline and notes that the numbers “might differ from other published numbers due to slightly different choices in the evaluation”. This site has not run them. The comparison models are the ones Microsoft chose in May 2024.
| Benchmark | Phi-3 Small 8K | Llama 3 8B Instruct | Gemma 7B | Mixtral 8x7B | GPT-3.5 Turbo 1106 |
|---|---|---|---|---|---|
| MMLU (5-shot) | 75.7 | 66.5 | 63.6 | 70.5 | 71.4 |
| HumanEval (0-shot) | 61.0 | 60.4 | 34.1 | 37.8 | 62.2 |
| MBPP (3-shot) | 71.7 | 67.7 | 51.5 | 60.2 | 77.8 |
| GSM8K, chain of thought (8-shot) | 89.6 | 77.4 | 59.8 | 64.7 | 78.1 |
| BigBench Hard (3-shot) | 79.1 | 51.5 | 59.6 | 69.7 | 68.3 |
| ARC Challenge (10-shot) | 90.7 | 82.8 | 78.3 | 87.3 | 87.4 |
| TriviaQA (5-shot) | 58.1 | 67.7 | 72.3 | 82.2 | 85.8 |
| Average of 19 benchmarks | 75.7 | 69.4 | 61.8 | 69.8 | 74.3 |
The TriviaQA row is the honest weak spot: 58.1, below every other model in the table. Microsoft's own category summary puts factual knowledge at 38.6. Phi-3 Small reasons well for its size and knows less than larger models, which is what you would expect from a model trained on a filtered, reasoning-heavy dataset.
What replaced Phi-3 Small
Microsoft did not release a 7B model in the Phi-3.5 or Phi-4 generations. The line continued at 3.8B and 14B, and both of those run on Ollama. Sizes are from the Ollama tag lists linked under Sources.
| Model | Ollama tag | Download | Note |
|---|---|---|---|
| Phi-4 (14B) | phi4 | 9.1 GB | Released 12 December 2024, MIT, 16K context. Fits a 12 GB card. |
| Phi-4 Mini (3.8B) | phi4-mini | 2.5 GB | The current small Phi. |
| Phi-3.5 Mini (3.8B) | phi3.5 | 2.2 GB | The refresh of Phi-3 Mini. |
| Phi-3 Medium (14B) | phi3:medium | 7.9 GB | The Phi-3 size above Small that Ollama does carry. |
For 7B to 8B models from other vendors that do run on Ollama, see Llama 3.1 8B, Qwen 2.5 7B, Mistral 7B and Gemma 7B, or go by card with the best LLM for 8 GB of VRAM.
Frequently asked questions
Can I run ollama pull phi3:small?
No. The tag does not exist. Ollama's phi3 library has 3.8b (Mini) and 14b (Medium) tags only. An earlier version of this page showed that command; it was wrong.
How much RAM does phi3:mini need on Ollama?
About 5 GB free at a 4K context, so 8 GB of system RAM is enough. The download is 2.2 GB. Memory use grows with the context length you set, by roughly 0.4 GB per thousand tokens.
Is there a Phi-3 Small with a longer context?
Yes, Phi-3-small-128k-instruct. It has the same parameter count and the same runtime requirements as the 8K model.
Can I use Phi-3 Small commercially?
The model card states the MIT licence, which permits commercial use. Phi-4 is MIT as well.
Is Phi-3 Small still worth running?
In my view only if you already run ONNX Runtime or need this exact model. It takes more setup than anything on Ollama, and Phi-4 is newer and a one-line install. If 8 GB of VRAM is your limit, Phi-4 Mini or Phi-3.5 Mini are the simpler choice.
Sources
- Phi-3-small-8k-instruct model card — Microsoft: specifications, hardware notes, benchmark table, licence
- Phi-3 Technical Report — arXiv:2404.14219, 22 April 2024
- Phi-3 Small ONNX builds — int4 and fp16 file sizes
- Phi-3 Cookbook — Microsoft's worked examples
- Ollama phi3 tags, phi4 tags and phi4-mini tags — every size quoted for Mini, Medium and Phi-4
- Phi-4 model card — release date, context length, licence
Ready to Go Beyond Tutorials?
25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Written by the Local AI Master Team
The team behind Local AI Master
We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.
Related Guides
Continue your local AI journey with these comprehensive guides
- PILLARLocal AI Models Directory: Every Model Compared
- Alpaca 7B: Stanford\
- Amazon Chronos: Time Series Forecasting Models (Complete Guide)
- Anima 2.9B on 8GB: The Anime Model Taking Over Civitai
- Aquila 7B by BAAI: Chinese-English Bilingual (FlagAI)
- Baichuan2-13B: Chinese LLM | 59% CMMLU, Bilingual, Free License 2026
- Bark by Suno AI: Open-Source Text-to-Audio Generation Guide
- ChatGLM3-6B: Tsinghua Chinese AI | Code Interpreter, 6GB RAM 2026
- Claude 3 Opus Review: Benchmarks, Pricing & API Guide 2026
- Claude 3 Sonnet Review: Benchmarks, API Pricing & Alternatives 2026
Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide
No spam. Unsubscribe with one click.
Found your model? Now build something with it.
25 hands-on courses — RAG, agents, fine-tuning — all running locally. First chapter free, no card.