★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once

Phi-3 Small 7B: sizes, requirements and why it is not on Ollama

Updated: September 29, 2026

Phi-3 Small is Microsoft's 7.4 billion parameter model from May 2024, and it is the one Phi-3 size that Ollama does not carry. There is no phi3:small tag. The weights are 14.8 GB at bf16 and the int4 ONNX build is 5.0 GB, so it runs on a 24 GB card with Transformers or on an 8 GB card with ONNX Runtime.

If you came here for Ollama sizes: phi3:mini is a 2.2 GB download and runs on a machine with 8 GB of RAM, and phi3:medium is 7.9 GB. Both have been replaced by Phi-4 and Phi-4 Mini, which are on Ollama.

Phi-3 Small hardware requirements

A 24 GB NVIDIA card for the original weights, or an 8 GB card for the int4 ONNX build. The arithmetic is parameters × bits per weight: 7.39 billion × 16 bits ÷ 8 is 14.8 GB, which is the size of the safetensors files on Hugging Face. The 5.0 GB int4 build works out to 5.4 bits per weight.

BuildSizeRuntimeWhat fits it
Original weights (bf16 safetensors)14.8 GBTransformers with trust_remote_codeA 24 GB card. A 16 GB card holds the weights with about 1 GB left for context.
ONNX fp16 (CUDA)15.6 GBONNX RuntimeA 24 GB card.
ONNX int4 (CUDA)5.0 GBONNX RuntimeAn 8 GB card.

There is a catch with the original weights. The model card says Phi-3 Small uses flash attention 2 and Triton blocksparse attention by default, “which requires certain types of GPU hardware to run”, and the only GPU types Microsoft lists are the NVIDIA A100, A6000 and H100. Microsoft's sample code asserts that CUDA is available. For anything else, including CPUs and Windows machines with AMD or Intel graphics, the card points to the ONNX builds.

Sizes are the folder totals reported by the Hugging Face API for microsoft/Phi-3-small-8k-instruct-onnx-cuda. The “what fits” column is arithmetic, not a measurement. To size other models, use the VRAM calculator.

Phi-3 model sizes on Ollama, and RAM for phi3:mini

phi3:mini is 2.2 GB and needs about 5 GB of free memory at a 4K context, so it runs on a machine with 8 GB of RAM or a 6 GB graphics card. phi3:medium is 7.9 GB. Ollama's library describes Phi-3 as Mini and Medium only; Small is absent.

ModelOllama tagDownloadNote
Phi-3 Mini (3.8B)phi3:mini2.2 GBSame file as phi3 and phi3:3.8b. It is the q4_0 of the 128K variant.
Phi-3 Mini 4Kphi3:mini-4k2.4 GBQ4_K_M of the 4K variant.
Phi-3 Mini, Q8_0phi3:3.8b-mini-4k-instruct-q8_04.1 GB
Phi-3 Mini, fp16phi3:3.8b-mini-4k-instruct-fp167.6 GB
Phi-3 Small (7B)nonen/aOllama does not publish it.
Phi-3 Medium (14B)phi3:medium7.9 GBSame file as phi3:14b.
Phi-3 Medium, Q8_0phi3:14b-medium-4k-instruct-q8_015 GB
Phi-3 Medium, fp16phi3:14b-medium-4k-instruct-fp1628 GB

The 5 GB figure for phi3:mini is the 2.2 GB file, plus 1.6 GB of KV cache for 4,096 tokens, plus a 1 GB allowance for buffers. Phi-3 Mini does not use grouped-query attention, so its cache is large for a model this size: 32 layers × 3072 × 2 × 2 bytes is 393 KB per token. That matters because Ollama lists phi3:mini with a 128K context. Filling all of it would take 51.5 GB of cache, so set the context length to what you need.

All sizes are from Ollama's phi3 tag list. The individual pages for Phi-3 Mini 3.8B and Phi-3 Medium 14B go further, and the Ollama system requirements guide covers memory by model size.

Specifications

From the model card and the config.json. The family is described in the Phi-3 Technical Report (arXiv:2404.14219, April 2024), which covers the 3.8B, 7B and 14B models together.

FieldPhi-3 Small
Parameters7,392,274,432 (7.4B)
ArchitectureDense decoder-only transformer with alternating dense and blocksparse attention
Layers32
Hidden size4096
Attention heads32, with 8 key/value heads
Vocabulary100,352 (tiktoken)
Context length8K, with a separate 128K variant
Training data4.8 trillion tokens, cutoff October 2023
Training run1024 H100-80G GPUs for 18 days
Weights released21 May 2024
LicenceMIT

Two of these set Small apart from its siblings. It uses the tiktoken tokenizer with a 100,352-token vocabulary, where Mini and Medium use a 32,064-token vocabulary. And it alternates dense and blocksparse attention layers, where the other two are standard transformers. Its config declares its own model type, phi3small. That is the reason the usual GGUF route did not pick it up: a llama.cpp request for conversion support, issue 8241, was closed as not planned in July 2024.

How to run Phi-3 Small

With Transformers on a supported NVIDIA GPU. The model card asks for tiktoken and triton and for trust_remote_code=True, because the model code ships in the repository. This is Microsoft's sample, shortened:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline

model_id = "microsoft/Phi-3-small-8k-instruct"
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    trust_remote_code=True,
)
assert torch.cuda.is_available(), "This model needs a GPU to run ..."
device = torch.cuda.current_device()
model = model.to(device)
tokenizer = AutoTokenizer.from_pretrained(model_id)

pipe = pipeline("text-generation", model=model, tokenizer=tokenizer, device=device)
messages = [{"role": "user", "content": "What about solving an 2x + 3 = 7 equation?"}]
output = pipe(messages, max_new_tokens=500, return_full_text=False, do_sample=False)
print(output[0]["generated_text"])

On other hardware, use the ONNX builds with ONNX Runtime GenAI. Microsoft publishes int4 builds for DirectML, CUDA, and CPU and mobile, plus an fp16 CUDA build. If what you want is a small Microsoft model behind ollama run, that is not this model; pull one of the models in the table below.

Benchmarks Microsoft reported

Copied from the Phi-3-small-8k-instruct model card. Microsoft ran every column through the same internal pipeline and notes that the numbers “might differ from other published numbers due to slightly different choices in the evaluation”. This site has not run them. The comparison models are the ones Microsoft chose in May 2024.

BenchmarkPhi-3 Small 8KLlama 3 8B InstructGemma 7BMixtral 8x7BGPT-3.5 Turbo 1106
MMLU (5-shot)75.766.563.670.571.4
HumanEval (0-shot)61.060.434.137.862.2
MBPP (3-shot)71.767.751.560.277.8
GSM8K, chain of thought (8-shot)89.677.459.864.778.1
BigBench Hard (3-shot)79.151.559.669.768.3
ARC Challenge (10-shot)90.782.878.387.387.4
TriviaQA (5-shot)58.167.772.382.285.8
Average of 19 benchmarks75.769.461.869.874.3

The TriviaQA row is the honest weak spot: 58.1, below every other model in the table. Microsoft's own category summary puts factual knowledge at 38.6. Phi-3 Small reasons well for its size and knows less than larger models, which is what you would expect from a model trained on a filtered, reasoning-heavy dataset.

What replaced Phi-3 Small

Microsoft did not release a 7B model in the Phi-3.5 or Phi-4 generations. The line continued at 3.8B and 14B, and both of those run on Ollama. Sizes are from the Ollama tag lists linked under Sources.

ModelOllama tagDownloadNote
Phi-4 (14B)phi49.1 GBReleased 12 December 2024, MIT, 16K context. Fits a 12 GB card.
Phi-4 Mini (3.8B)phi4-mini2.5 GBThe current small Phi.
Phi-3.5 Mini (3.8B)phi3.52.2 GBThe refresh of Phi-3 Mini.
Phi-3 Medium (14B)phi3:medium7.9 GBThe Phi-3 size above Small that Ollama does carry.

For 7B to 8B models from other vendors that do run on Ollama, see Llama 3.1 8B, Qwen 2.5 7B, Mistral 7B and Gemma 7B, or go by card with the best LLM for 8 GB of VRAM.

Frequently asked questions

Can I run ollama pull phi3:small?

No. The tag does not exist. Ollama's phi3 library has 3.8b (Mini) and 14b (Medium) tags only. An earlier version of this page showed that command; it was wrong.

How much RAM does phi3:mini need on Ollama?

About 5 GB free at a 4K context, so 8 GB of system RAM is enough. The download is 2.2 GB. Memory use grows with the context length you set, by roughly 0.4 GB per thousand tokens.

Is there a Phi-3 Small with a longer context?

Yes, Phi-3-small-128k-instruct. It has the same parameter count and the same runtime requirements as the 8K model.

Can I use Phi-3 Small commercially?

The model card states the MIT licence, which permits commercial use. Phi-4 is MIT as well.

Is Phi-3 Small still worth running?

In my view only if you already run ONNX Runtime or need this exact model. It takes more setup than anything on Ollama, and Phi-4 is newer and a one-line install. If 8 GB of VRAM is your limit, Phi-4 Mini or Phi-3.5 Mini are the simpler choice.

Sources

Reading now
Join the discussion

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📅 Published: 2024-04-23🔄 Last Updated: September 29, 2026✓ Manually Reviewed

Related Guides

Continue your local AI journey with these comprehensive guides

More on AI Models Directory
See the full AI Models Directory guide.
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Found your model? Now build something with it.

25 hands-on courses — RAG, agents, fine-tuning — all running locally. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators