★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Hardware Reference

Best Ollama Models for 8GB, 12GB, 16GB & 24GB VRAM (Full Table)

April 11, 2026
19 min read
Local AI Master Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Ollama’s running. Here’s what to build with it. Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.

Start free
Or own it for life — Lifetime $149, pay once

Published on April 11, 2026 · Updated September 2026 · 19 min read

Quick answer: the best Ollama models by GPU are Llama 3.1 8B for 8GB VRAM (Qwen 2.5 Coder 7B if you code), Qwen 2.5 14B for 12GB, Gemma 3 27B for 16GB, and Qwen 2.5 32B — or the much faster MoE Qwen 3 30B-A3B — for 24GB. Here is the RAM/VRAM each popular model needs at Q4_K_M (the default quant), plus the minimum GPU/unified memory to load it fully on-GPU. Full per-family tables and bandwidth-derived throughput ceilings are further down.

ModelQ4_K_M sizeMin VRAMFits on
Llama 3.2 1B0.8GB2GBAny GPU / phone-class
Llama 3.2 3B1.9GB3GB4GB GPU
Llama 3.1 8B4.9GB6GB8GB GPU
Qwen 2.5 7B4.4GB5.5GB8GB GPU
Qwen 2.5 Coder 7B4.4GB5.5GB8GB GPU
Gemma 3 12B7.3GB8.5GB12GB GPU
Qwen 2.5 14B8.7GB10GB12GB GPU
Gemma 3 27B15.9GB17GB16GB GPU (tight) / 24GB
Qwen 2.5 32B18.8GB20GB24GB GPU
Qwen 2.5 Coder 32B18.8GB20GB24GB GPU
Llama 3.3 70B40.0GB42GB48GB+ unified / dual 24GB
Llama 3.1 405B229.0GB232GBMulti-GPU server

Rule of thumb: min VRAM ≈ the Q4_K_M file size + ~1–1.5GB overhead (KV cache at 2K context). Best pick for 6GB: a 3–4B at Q4 (laptop GPUs); 8GB: Llama 3.1 8B (general) or Qwen 2.5 Coder 7B (code); 12GB: a 14B at Q4; 16GB: Gemma 3 27B Q4_K_M; 24GB: Qwen 2.5 32B Q4_K_M; 32GB: a 32B at Q5/Q6 with full context on an RTX 5090. Each of those links ranks the full shortlist for that card. RAM vs VRAM: models run fastest fully in VRAM; if they spill to system RAM, expect 2–10× slower generation.

Want the math done for a specific model and context length? Plug your numbers into our interactive VRAM calculator. For the full picture on CPU, disk, and OS-level needs beyond raw model size, see the Ollama system requirements guide, and for task-by-task picks rather than raw sizing, see best Ollama models. If you'd rather have hardware matched to a model for you than work it out from a table, the local model picker course walks through that decision step by step.


This page exists so you never have to google "how much VRAM does Llama 70B need" again. It collects the published Q4/Q5/FP16 file size for every popular Ollama model, the minimum memory to load it fully on-GPU, and the throughput ceiling that each model's size imposes on three common machines.

On the throughput numbers: we do not own an RTX 3060, an RTX 4090 or an M4 Max, and none of the figures below were timed on one. Every tok/s column is an arithmetic ceiling — published memory bandwidth divided by the model's weight size — and is labelled as such. That is a genuinely useful number, because it is the speed the hardware physically cannot exceed. It is not a prediction of what you will see.

Bookmark this page. Sizes and ceilings both move only when a model or a card changes, so it stays correct far longer than a benchmark run would.


How to Read This Table

Each model entry includes:

  • Q4_K_M size — 4-bit quantization, the sweet spot for most users (minimal quality loss, biggest memory savings)
  • Q5_K_M size — 5-bit quantization, slightly better quality, ~20% more memory
  • FP16 size — Full precision, maximum quality, 2x memory vs Q4
  • Min VRAM — The minimum VRAM needed to load the Q4_K_M version entirely on GPU (includes ~1GB overhead for KV cache at 2K context)
  • 3060 ceiling — Arithmetic throughput ceiling on an RTX 3060 12GB (entry-level AI GPU)
  • 4090 ceiling — Arithmetic throughput ceiling on an RTX 4090 24GB (high-end consumer)
  • M4 Max ceiling — Arithmetic throughput ceiling on an Apple M4 Max with 48GB unified memory

Where the ceiling numbers come from

Generating one token requires reading the entire weight set out of memory exactly once. That makes decoding memory-bandwidth bound, and it gives a hard upper limit that needs no benchmarking at all:

Ceiling (tok/s) = memory bandwidth (GB/s) ÷ model file size (GB)

The three reference machines, using each vendor's published bandwidth: RTX 3060 12GB — 360 GB/s, RTX 4090 — 1008 GB/s, Apple M4 Max — 546 GB/s. A dash means the model does not fit that machine's memory, so it would offload to CPU and the ceiling no longer applies.

Three caveats, and they matter:

  1. A ceiling is not a measurement. Real output lands below it — often well below — once sampling, attention over a growing KV cache, and runtime overhead are counted. Use the ceiling to rule things out, not to predict a number.
  2. At roughly 2GB of weights or less the ceiling stops being the binding constraint (marked ‡ in the tables). The four-digit figures on the 1B rows are arithmetically correct and practically meaningless: at that size, kernel launch overhead and sampling dominate long before bandwidth does. Read those rows as "bandwidth is not your problem here".
  3. MoE models break the formula (marked † in the tables). A Mixture-of-Experts model only activates a fraction of its weights per token, so bytes moved per token are far below the file size. Those rows carry an explanation rather than a number.

Because the ceiling depends only on published bandwidth and published file sizes, it does not drift between Ollama point releases the way a benchmark does. Ollama's latest stable release as of this September 2026 refresh is v0.34.0 — see the Ollama version history for what changed, or the Ollama system requirements guide for install sizes and hardware floors.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

Llama 3.x Family

The workhorse family. Meta's Llama models are the most-used open-weight models for good reason — strong quality across the board.

ModelParamsQ4_K_MQ5_K_MFP16Min VRAM3060 ceiling4090 ceilingM4 Max ceiling
Llama 3.2 1B1.24B0.8GB0.9GB2.5GB2GB450‡1260‡680‡
Llama 3.2 3B3.21B1.9GB2.3GB6.4GB3GB190‡530‡290‡
Llama 3.1 8B8.03B4.9GB5.7GB16.1GB6GB73205110
Llama 3.3 70B70.6B40.0GB48.0GB141.0GB42GB——14
Llama 3.1 405B405B229.0GB275.0GB810.0GB232GB———

‡ At roughly 2GB of weights or less, bandwidth stops being the binding constraint — see caveat 2 above.

Notes:

  • Llama 3.1 8B is the single most popular Ollama model. It fits on any 8GB GPU with room to spare.
  • Llama 3.3 70B at Q4_K_M needs 42GB — runs on M4 Max 48GB or dual 24GB GPUs. Does not fit on a single RTX 4090.
  • Llama 3.1 405B is impractical for consumer hardware. Included for completeness. Requires 4x A100 80GB or equivalent.
# Pull Llama models
ollama pull llama3.2:1b
ollama pull llama3.2:3b
ollama pull llama3.2        # defaults to 8B Q4_K_M
ollama pull llama3.3:70b-instruct-q4_K_M

How Much VRAM Does Llama 4 Need?

Meta's Llama 4 generation is built on Mixture-of-Experts (MoE), so the memory math is different from the dense Llama 3.x models above. You still have to fit every expert weight in memory (so the file is large), but only a fraction of parameters activate per token (so it runs faster than a dense model of the same total size). This is the single most common point of confusion people hit when sizing hardware for Llama 4.

ModelTotal / Active paramsQ4_K_MMin VRAMFits on
Llama 4 Scout109B / 17B active67GB (Ollama library)~68GB80GB+ unified / multi-GPU 80GB+ combined / H100 80GB
Llama 4 Maverick400B / 17B active~210–230GB~230GBMulti-GPU server only

Notes:

  • Llama 4 Scout is the smallest Llama 4 and still needs 67GB at Q4_K_M (per the Ollama library) because all 16 experts must be resident. It does not fit a single 24GB consumer GPU — plan on an 80GB+ Apple Silicon machine, a multi-GPU rig with 80GB+ combined VRAM (dual RTX 5090s' 64GB combined is not quite enough), or a data-center card. Once loaded, it generates at the speed of a ~17B dense model thanks to MoE routing.
  • Llama 4 Maverick is firmly server-class. Treat it like the 405B entry in the table: included for completeness, not for a desktop.
  • If Llama 4 is out of reach, a Q4_K_M Llama 3.3 70B (40GB) on 48GB Apple Silicon remains the most accessible top-tier Llama for local use. Verify the exact tag and size against the official Ollama library before pulling, since MoE quant sizes vary by release.

Qwen 2.5 / Qwen 3 Family

Alibaba's Qwen models punch above their weight on code, math, and multilingual tasks. Qwen 2.5 Coder is the best local coding model at every size.

ModelParamsQ4_K_MQ5_K_MFP16Min VRAM3060 ceiling4090 ceilingM4 Max ceiling
Qwen 2.5 0.5B0.49B0.4GB0.5GB1.0GB1.5GB900‡2520‡1370‡
Qwen 2.5 1.5B1.54B1.0GB1.2GB3.1GB2GB360‡1010‡550‡
Qwen 2.5 3B3.09B1.9GB2.2GB6.2GB3GB190‡530‡290‡
Qwen 2.5 7B7.62B4.4GB5.2GB15.2GB5.5GB82230125
Qwen 2.5 14B14.8B8.7GB10.3GB29.5GB10GB4111563
Qwen 2.5 32B32.5B18.8GB22.5GB65.0GB20GB—5429
Qwen 2.5 72B72.7B42.0GB50.0GB145.0GB44GB——13
Qwen 3 8B8.2B5.0GB5.9GB16.4GB6.5GB72200110
Qwen 3 14B14.8B~9.0GB10.6GB29.6GB10.5GB4011061
Qwen 3 30B-A3B (MoE)30.5B / 3.3B active19GB—61.0GB~20GB—MoE†MoE†
Qwen 3 32B32.8B19.2GB23.0GB65.6GB21GB—5228

‡ At roughly 2GB of weights or less, bandwidth stops being the binding constraint — see caveat 2 above. † MoE: only ~3.3B of 30.5B parameters activate per token, so bytes moved per token are closer to a 2GB model than a 17GB one. The formula does not apply; expect throughput in the same neighbourhood as a small dense model.

Notes:

  • Qwen 2.5 7B is the go-to if you want code+math strength on an 8GB GPU.
  • Qwen 2.5 14B Q4_K_M at 8.7GB barely fits on a 12GB RTX 3060 — you get it running but context window is limited.
  • Qwen 2.5 32B is the sweet spot for 24GB GPUs. Fits with room for a decent context window.
  • Qwen 3 models use a hybrid thinking architecture (think/no-think modes). Slightly larger than Qwen 2.5 at the same parameter count.
  • Qwen 3 30B-A3B is the standout for 24GB GPUs in 2026. It is a Mixture-of-Experts model: ~30B total params (Ollama's default qwen3:30b-a3b / qwen3:30b tag is 19GB, needing ~20GB to load — verified against the Ollama library September 2026), but only ~3.3B activate per token. That is the whole point — you pay 32B-class memory but move roughly a tenth of those bytes per token, so it generates like a small dense model rather than a 30B one. If you have the VRAM, it is often a better pick than the dense Qwen 3 32B for interactive use.
  • Qwen3.6 27B (ollama pull qwen3.6:27b) is the newest dense pick for 24GB cards — 18GB per the Ollama library, which against the RTX 4090's 1008 GB/s puts the arithmetic ceiling near 56 tok/s. The Qwen3.6-27B guide has full specs, quant sizes and VRAM figures. If you want the whole Ollama workflow — install, model library, Modelfiles, GPU offload, serving — in one place, the Ollama Mastery course covers it; the first chapter is free with an account.
# Pull Qwen models
ollama pull qwen2.5:0.5b
ollama pull qwen2.5:7b
ollama pull qwen2.5:14b
ollama pull qwen2.5:32b
ollama pull qwen2.5:72b
ollama pull qwen3:8b
ollama pull qwen3:14b
ollama pull qwen3:30b-a3b      # MoE, 19GB, fits 24GB
ollama pull qwen3:32b

Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

Gemma 3 & Gemma 4 Family

Google's Gemma 3 models are surprisingly strong at small sizes. Gemma 3 4B punches way above its weight and is a top pick for constrained devices. The Gemma 4 generation landed in summer 2026 — Gemma 3 with published file sizes first, then the Gemma 4 planning estimates.

ModelParamsQ4_K_MQ5_K_MFP16Min VRAM3060 ceiling4090 ceilingM4 Max ceiling
Gemma 3 1B1.0B0.7GB0.8GB2.0GB2GB510‡1440‡780‡
Gemma 3 4B3.9B2.5GB3.0GB7.8GB3.5GB145400220
Gemma 3 12B12.2B7.3GB8.7GB24.4GB8.5GB4914075
Gemma 3 27B27.2B15.9GB19.0GB54.4GB17GB—6334

‡ At roughly 2GB of weights or less, bandwidth stops being the binding constraint — see caveat 2 above.

Notes:

  • Gemma 3 4B at 2.5GB Q4_K_M is excellent for Raspberry Pi 5 (8GB) or old laptops.
  • Gemma 3 12B is a strong 12GB GPU choice, rivaling models twice its size on instruction following.
  • Gemma 3 27B fits on a single RTX 4090 (24GB) at Q4_K_M with tight context, or comfortably on 32GB Apple Silicon.
# Pull Gemma models
ollama pull gemma3:1b
ollama pull gemma3:4b
ollama pull gemma3:12b
ollama pull gemma3:27b

Gemma 4 (Summer 2026)

Gemma 4 has since shipped on Ollama with published sizes: the 12B Unified is a 7.6GB download, so a 12GB card runs it comfortably, and both the 26B MoE and the 31B dense flagship fit a single 24GB card. The whole family ships under Apache 2.0. Sizes below are the exact download sizes from the Ollama library (checked September 2026), not estimates — no ceiling column because per-chip bandwidth still varies too much to generalize.

ModelParamsSize (Ollama default tag)Min VRAMFits on
Gemma 4 12B Unified12B (text+image, 256K ctx)7.6GB~9GB12GB GPU (8GB only with short context)
Gemma 4 26B26B MoE, ~3.8B active19GB~20GB24GB GPU
Gemma 4 31B31B dense20GB~21GB24GB GPU

Notes:

  • The E2B and E4B members are not as tiny as their names suggest: Ollama's default gemma4:e2b tag is 7.2GB and gemma4:e4b is 9.6GB, because the effective-parameter count excludes the multimodal embeddings that still have to be resident. Plan on an 8GB+ GPU or 16GB Mac for either, not a phone-class device — see the Ollama system requirements guide for the full 2026-model memory table.
  • The 12B Unified handles text and images in one model with a 256K context. Long context costs VRAM on top of the weights, so treat 12-16GB as the practical floor if you actually want that window.
  • Check the current tags on the official Ollama library before pulling, and see our full Gemma 4 hardware guide for the tier-by-tier breakdown.

Phi-4 Family

Microsoft's Phi-4 models achieve remarkable reasoning for their size. Phi-4 3.8B consistently beats 7B models from other families on logic and math tasks.

ModelParamsQ4_K_MQ5_K_MFP16Min VRAM3060 ceiling4090 ceilingM4 Max ceiling
Phi-4 Mini (3.8B)3.82B2.3GB2.8GB7.6GB3.5GB155440235
Phi-4 (14B)14.0B8.2GB9.8GB28.0GB9.5GB4412567

Notes:

  • Phi-4 Mini at 2.3GB Q4_K_M is the best reasoning model you can fit on a 4GB GPU.
  • Phi-4 14B needs a 12GB GPU minimum. Performance is excellent for code review and analytical tasks.
# Pull Phi models
ollama pull phi4-mini       # 3.8B
ollama pull phi4            # 14B

Mistral / Mixtral Family

Mistral's models and their Mixture-of-Experts (MoE) Mixtral variants. MoE models use more disk space but activate only a fraction of parameters per token, giving better quality per FLOP.

ModelParamsQ4_K_MQ5_K_MFP16Min VRAM3060 ceiling4090 ceilingM4 Max ceiling
Mistral 7B v0.37.25B4.4GB5.1GB14.5GB5.5GB82230125
Mistral Small 24B24.0B14.0GB16.8GB48.0GB15.5GB—7239
Mixtral 8x7B46.7B (12.9B active)26.4GB31.7GB93.4GB28GB——MoE†
Mixtral 8x22B176B (39B active)80.0GB96.0GB352.0GB82GB———

† MoE: only 12.9B of 46.7B parameters activate per token, so roughly 7.3GB moves per token rather than the full 26.4GB. Against the M4 Max's 546 GB/s that is a ceiling near 75 tok/s rather than 21 — the file size overstates the work.

Notes:

  • Mistral 7B is a solid all-rounder, competitive with Llama 3.1 8B. Slightly smaller file size.
  • Mixtral 8x7B has 46.7B total params but only activates 12.9B per token. Needs 28GB VRAM to hold, but moves roughly a quarter of that per token, so it generates like a ~13B model. Quality approaches 70B dense models.
  • Mixtral 8x22B is a server-class model. 80GB minimum VRAM. Requires A100 80GB or multi-GPU.
# Pull Mistral/Mixtral models
ollama pull mistral          # 7B
ollama pull mistral-small    # 24B
ollama pull mixtral          # 8x7B

DeepSeek R1 Distills

DeepSeek's R1 reasoning model distilled into smaller architectures. These models "think" step-by-step and show their reasoning chain.

ModelBaseQ4_K_MQ5_K_MFP16Min VRAM3060 ceiling4090 ceilingM4 Max ceiling
DeepSeek R1 1.5BQwen 2.5 1.5B1.0GB1.2GB3.1GB2GB360‡1010‡550‡
DeepSeek R1 7BQwen 2.5 7B4.4GB5.2GB15.2GB5.5GB82230125
DeepSeek R1 8BLlama 3.1 8B4.9GB5.7GB16.1GB6GB73205110
DeepSeek R1 14BQwen 2.5 14B8.7GB10.3GB29.5GB10GB4111563
DeepSeek R1 32BQwen 2.5 32B18.8GB22.5GB65.0GB20GB—5429
DeepSeek R1 70BLlama 3.3 70B40.0GB48.0GB141.0GB42GB——14

‡ At roughly 2GB of weights or less, bandwidth stops being the binding constraint — see caveat 2 above.

Notes:

  • R1 distills have the same VRAM requirements and the same ceilings as their base models (same architecture, same parameter count, same file size).
  • The "thinking" tokens add output length, and the ceiling is per token — so a reasoning model at the same tok/s ceiling still takes longer to reach an answer. A simple question might generate 500+ reasoning tokens before the final answer. Budget accordingly.
  • R1 7B on an 8GB GPU gives you step-by-step reasoning that rivals much larger non-reasoning models on math and logic.
# Pull DeepSeek R1 distills
ollama pull deepseek-r1:1.5b
ollama pull deepseek-r1:7b
ollama pull deepseek-r1:8b
ollama pull deepseek-r1:14b
ollama pull deepseek-r1:32b
ollama pull deepseek-r1:70b

Code Models

Specialized models for code generation, completion, and review. These are fine-tuned on code and perform better than general-purpose models on programming tasks.

ModelParamsQ4_K_MQ5_K_MFP16Min VRAM3060 ceiling4090 ceilingM4 Max ceiling
Qwen 2.5 Coder 1.5B1.54B1.0GB1.2GB3.1GB2GB360‡1010‡550‡
Qwen 2.5 Coder 7B7.62B4.4GB5.2GB15.2GB5.5GB82230125
Qwen 2.5 Coder 14B14.8B8.7GB10.3GB29.5GB10GB4111563
Qwen 2.5 Coder 32B32.5B18.8GB22.5GB65.0GB20GB—5429
Qwen3-Coder 30B (MoE)30.5B / ~3B active19GB——~20GB—MoE†—
CodeLlama 7B6.74B3.8GB4.5GB13.5GB5GB95265145
CodeLlama 13B13.0B7.4GB8.8GB26.0GB8.5GB4913574
CodeLlama 34B33.7B19.1GB22.8GB67.4GB20.5GB—5329
StarCoder2 3B3.03B1.8GB2.2GB6.1GB3GB200‡560‡305‡
StarCoder2 7B6.74B3.8GB4.5GB13.5GB5GB95265145
StarCoder2 15B15.5B9.0GB10.8GB31.0GB10.5GB4011061

‡ At roughly 2GB of weights or less, bandwidth stops being the binding constraint — see caveat 2 above. † MoE: only ~3B of 30.5B parameters activate per token, so roughly 1.9GB moves per token rather than the full 19GB. The formula does not apply — expect small-dense-model throughput from a 32B-class model.

Notes:

  • The Qwen 3 wave changed the coding pecking order. Qwen3-Coder 30B is an MoE with only ~3B params active per token, so it delivers 32B-class code while moving roughly a tenth of a dense 32B's bytes per token — that is where the speed advantage comes from, not from a smaller file. It also brings a 256K native context built for agent workflows. Qwen 2.5 Coder 32B is still the strongest dense coder and a fine choice if you already run it, but the MoE is the one to pull first. Full rankings in our best local coding models guide.
  • Do not undersize Qwen3-Coder-Next (80B total, ~3B active) — it is a bigger MoE than it sounds. Ollama's default qwen3-coder-next tag is Q4_K_M at 52GB, so despite the small active-parameter count it needs a 64GB-class machine (dual GPU or a high-memory Apple Silicon Mac) to hold the full expert set resident, not a 16GB card. Check the current size on the Ollama library before planning hardware around it.
  • Qwen 2.5 Coder 7B is the best coding model for 8GB GPUs. Outperforms CodeLlama 13B despite being smaller.
  • CodeLlama is aging but still widely used. If you are already using it, consider switching to Qwen 2.5 Coder at the same size.
  • StarCoder2 excels at code completion (fill-in-the-middle) rather than instruction following.
# Pull code models
ollama pull qwen3-coder:30b     # MoE, 19GB, fits 24GB
ollama pull qwen3-coder-next    # 3B active, runs on 16GB
ollama pull qwen2.5-coder:7b
ollama pull qwen2.5-coder:14b
ollama pull qwen2.5-coder:32b
ollama pull codellama:7b
ollama pull codellama:13b
ollama pull codellama:34b
ollama pull starcoder2:3b
ollama pull starcoder2:7b
ollama pull starcoder2:15b

Legacy Models: Llama 2 Memory Requirements

Llama 2 7B at FP16 needs ~14GB of VRAM (13.5GB of weights plus KV cache); at Q4_K_M it drops to ~4.1GB and runs on any 6GB GPU. The 13B and 70B follow the same math: two bytes per parameter at FP16, roughly 30% of that at Q4_K_M.

ModelParamsQ4_K_MFP16Min VRAM (Q4_K_M)Min VRAM (FP16)
Llama 2 7B6.74B4.1GB13.5GB5GB~14GB
Llama 2 13B13.0B7.9GB26.0GB9GB~27GB
Llama 2 70B69B41.4GB138.0GB43GB~140GB

Notes:

  • Running Llama 2 7B unquantized (FP16) means a 16GB card — it does not fit 12GB once the KV cache is counted. In practice almost nobody should: Q4_K_M carries the usual 1-3% quality cost at less than a third of the memory.
  • Ollama's default llama2 tags ship Q4_0, slightly smaller than Q4_K_M: 3.8GB for 7B, 7.4GB for 13B, 39GB for 70B.
  • Llama 2 is two generations old. Unless you are reproducing older research or maintaining an existing pipeline, Llama 3.1 8B beats Llama 2 13B across the board at a third of the memory.
# Pull Llama 2 (legacy)
ollama pull llama2          # 7B, Q4_0
ollama pull llama2:13b
ollama pull llama2:70b

Quick Reference: What Fits on Your GPU

Find your GPU (or unified memory) capacity and see what models you can run at full speed (100% VRAM, no CPU offload).

These tables answer "what fits". For "how fast", divide your card's bandwidth by the size in the third column — the ceiling formula from the "How to Read This Table" section above. Handy reference bandwidths: RTX 4060 272 GB/s, RTX 3060 12GB 360 GB/s, RTX 4070 504 GB/s, RTX 3090 936 GB/s, RTX 4090 1008 GB/s, RTX 5090 1792 GB/s. Apple publishes per-chip figures that vary widely across the M-series, so check the one you own.

8GB VRAM (RTX 3060 8GB, RTX 4060, M1/M2 8GB)

ModelQuantSizeQuality
Llama 3.1 8BQ4_K_M4.9GBStrong general
Qwen 2.5 7BQ4_K_M4.4GBBest for code/math
Gemma 3 4BQ5_K_M3.0GBGreat for small tasks
Phi-4 Mini 3.8BQ5_K_M2.8GBBest tiny reasoner
DeepSeek R1 7BQ4_K_M4.4GBChain-of-thought
Qwen 2.5 Coder 7BQ4_K_M4.4GBBest coding for 8GB

Top pick: Llama 3.1 8B Q4_K_M for general use, Qwen 2.5 Coder 7B for programming. On an RTX 4060 (272 GB/s) that 4.9GB file puts the ceiling near 55 tok/s.

12GB VRAM (RTX 3060 12GB, RTX 4070)

Everything from 8GB, plus:

ModelQuantSizeQuality
Qwen 2.5 14BQ4_K_M8.7GBSignificant quality jump
Gemma 3 12BQ4_K_M7.3GBExcellent instruction
Phi-4 14BQ4_K_M8.2GBStrong reasoning
Llama 3.1 8BQ5_K_M5.7GBHigher quality 8B
CodeLlama 13BQ4_K_M7.4GBSolid code model

Top pick: Qwen 2.5 14B Q4_K_M. The jump from 7B to 14B is the single biggest quality improvement per dollar in local AI. On an RTX 3060 12GB (360 GB/s) the 8.7GB file caps generation at about 41 tok/s; an RTX 4070 (504 GB/s) raises that to about 58. If you are still choosing a card rather than sizing one you own, our RTX 5070 for local AI breakdown covers exactly what 12GB runs and when paying up for 16GB is worth it.

16GB VRAM (RTX 4080, RTX 5060 Ti, M1 Pro/M2 Pro 16GB)

Everything from 12GB, plus:

ModelQuantSizeQuality
Mistral Small 24BQ4_K_M14.0GBStrong all-rounder
Gemma 3 27BQ4_K_M15.9GBNear-70B quality
Qwen 2.5 14BQ5_K_M10.3GBHigher quality 14B

Top pick: Gemma 3 27B Q4_K_M squeezes in and delivers impressive quality. Note that at 15.9GB it leaves almost nothing for the KV cache on a 16GB card, so plan on a short context or step down a size.

24GB VRAM (RTX 3090, RTX 4090, M3 Pro 24GB)

Everything from 16GB, plus:

ModelQuantSizeQuality
Qwen 2.5 32BQ4_K_M18.8GBNear-70B quality
Qwen3-Coder 30B (MoE)Q4_K_M19GBBest local coding
Qwen 2.5 Coder 32BQ4_K_M18.8GBStrongest dense coder
DeepSeek R1 32BQ4_K_M18.8GBBest local reasoning
CodeLlama 34BQ4_K_M19.1GBMature code model
Qwen 2.5 14BFP1629.5GB— (too large)

Top pick: Qwen 2.5 32B Q4_K_M. Best overall dense model that fits on a single consumer GPU. Its 18.8GB against an RTX 4090's 1008 GB/s is a ceiling near 54 tok/s, and near 50 on a 3090's 936 GB/s — the two cards are much closer here than their price gap suggests, because capacity is identical and bandwidth differs by only 8%.

32GB Unified Memory (M2 Pro/Max 32GB, M3 Pro 36GB)

Everything from 24GB, plus:

ModelQuantSizeQuality
Mixtral 8x7BQ4_K_M26.4GBMoE, broad knowledge
Qwen 2.5 32BQ5_K_M22.5GBHigher quality 32B

Apple's unified-memory bandwidth varies enormously across chips at the same capacity — a Pro and a Max with identical RAM are not the same machine for inference. Check your specific chip's published figure before applying the ceiling formula.

48GB+ Unified Memory (M3 Max 48GB, M4 Max 48GB+)

Everything from 32GB, plus:

ModelQuantSizeQuality
Llama 3.3 70BQ4_K_M40.0GBTop-tier open model
Qwen 2.5 72BQ4_K_M42.0GBExcellent multilingual
DeepSeek R1 70BQ4_K_M40.0GBBest open reasoning

Top pick: Llama 3.3 70B Q4_K_M. Running a 70B model on a laptop is genuinely impressive — but do the arithmetic before you expect it to feel snappy: 40GB against the M4 Max's 546 GB/s is a ceiling of about 14 tok/s, and that is the number you cannot beat, not the number you will get.

64GB+ (M4 Ultra, dual GPU, server)

Everything from 48GB, plus:

ModelQuantSizeQuality
Llama 3.3 70BQ5_K_M48.0GBBest quality 70B
Qwen 2.5 72BQ5_K_M50.0GBHigher quality 72B
Mixtral 8x22BQ4_K_M80.0GBNeeds 82GB+ (64GB not enough)

Stepping from Q4_K_M to Q5_K_M costs about 20% more bytes per token, so it costs about 20% off the throughput ceiling too. That is the real price of the quality bump at this size.

For more on choosing hardware for your target models, see our RAM requirements guide and VRAM requirements guide.


Ollama Pull Commands: Every Model

Copy-paste ready. Every model referenced in this article:

# === LLAMA FAMILY ===
ollama pull llama3.2:1b
ollama pull llama3.2:3b
ollama pull llama3.2                     # 8B, default quant
ollama pull llama3.3:70b-instruct-q4_K_M

# === QWEN FAMILY ===
ollama pull qwen2.5:0.5b
ollama pull qwen2.5:1.5b
ollama pull qwen2.5:3b
ollama pull qwen2.5:7b
ollama pull qwen2.5:14b
ollama pull qwen2.5:32b
ollama pull qwen2.5:72b
ollama pull qwen3:8b
ollama pull qwen3:32b

# === GEMMA FAMILY ===
ollama pull gemma3:1b
ollama pull gemma3:4b
ollama pull gemma3:12b
ollama pull gemma3:27b

# === PHI FAMILY ===
ollama pull phi4-mini
ollama pull phi4

# === MISTRAL/MIXTRAL ===
ollama pull mistral
ollama pull mistral-small
ollama pull mixtral

# === DEEPSEEK R1 DISTILLS ===
ollama pull deepseek-r1:1.5b
ollama pull deepseek-r1:7b
ollama pull deepseek-r1:8b
ollama pull deepseek-r1:14b
ollama pull deepseek-r1:32b
ollama pull deepseek-r1:70b

# === CODE MODELS ===
ollama pull qwen3-coder:30b
ollama pull qwen3-coder-next
ollama pull qwen2.5-coder:1.5b
ollama pull qwen2.5-coder:7b
ollama pull qwen2.5-coder:14b
ollama pull qwen2.5-coder:32b
ollama pull codellama:7b
ollama pull codellama:13b
ollama pull codellama:34b
ollama pull starcoder2:3b
ollama pull starcoder2:7b
ollama pull starcoder2:15b

# === LEGACY (LLAMA 2) ===
ollama pull llama2
ollama pull llama2:13b
ollama pull llama2:70b

Browse the full model library at ollama.com/library.


How Much Does Context Length Add to VRAM?

The "Min VRAM" column in every table above assumes a modest 2K-token context. In real use — long chats, RAG pipelines, agents, large code files — the KV cache grows with context length and can quietly become the thing that pushes you out of VRAM. This is the #1 reason a model that "should fit" suddenly spills to CPU. RAG setups pay a second tax: an embedder sits in VRAM alongside the chat model, though the best Ollama embedding models for RAG are small enough that a few hundred MB to ~1GB covers it.

The KV cache scales roughly linearly with context. A useful approximation for a typical dense model at Q4_K_M:

Extra VRAM for context ≈ 0.5 GB per 2K tokens (7-8B model)
                       ≈ 1.0 GB per 2K tokens (14B model)
                       ≈ 1.5-2.0 GB per 2K tokens (32B+ model)
Model (Q4_K_M)2K ctx8K ctx32K ctx
Llama 3.1 8B~6GB~7.5GB~14GB
Qwen 2.5 14B~10GB~13GB~22GB
Qwen 2.5 32B~20GB~24GB~36GB

Notice what this means in practice: a 32B model that fits a 24GB GPU at 2K context will not fit the same card at 32K context — it needs ~36GB. If you run long-context workloads, size your hardware against the right-hand column, not the headline number. Two ways to claw back memory: enable KV-cache quantization (OLLAMA_KV_CACHE_TYPE=q8_0 roughly halves cache size at a tiny quality cost), or set a smaller num_ctx for the model. These are approximate; exact KV-cache size depends on the model's head count and hidden dimension. The broader VRAM requirements guide covers bandwidth and bus-width effects on top of capacity.

If you are working with a tight memory budget, the safest move is to drop a size class and run a smaller model at a longer context rather than a bigger model that constantly spills — see the best local AI models for 8GB RAM for picks that leave headroom for context.


Memory Math: How to Calculate Any Model

If a model is not in this table, you can estimate its VRAM requirement:

FP16 size (GB)  = Parameters (B) × 2
Q4_K_M size     ≈ FP16 × 0.28 to 0.32  (varies by architecture)
Q5_K_M size     ≈ FP16 × 0.34 to 0.38
Q8_0 size       ≈ FP16 × 0.50 to 0.55

Min VRAM needed = Model file size + 1.0 GB (KV cache at 2K context)
                  + 0.5 GB per 2K additional context tokens

Example: A new 20B model you want to run at Q4_K_M:

  • FP16 size: 20 × 2 = 40GB
  • Q4_K_M: 40 × 0.30 = ~12GB
  • Min VRAM: 12 + 1.0 = ~13GB
  • Fits on a 16GB GPU with room for 4K context

For a thorough explanation of quantization levels and their quality tradeoffs, read our quantization explained guide. To find the best models for tight memory budgets, see best models for 8GB RAM.


Frequently Asked Questions

See the FAQ section below for answers to common questions about Ollama memory requirements.


Keep This Bookmarked

This table gets updated as new models release — the August 2026 refresh added Gemma 4, the Qwen 3 coder wave, and the Llama 2 legacy table; the September 2026 refresh re-verified every Gemma 4, Qwen3.6, Qwen3-Coder and Llama 4 size against the live Ollama library and corrected several that had drifted. The Ollama ecosystem moves fast — new model families appear every few weeks, and existing ones get updated quantization options.

The core principle stays constant: check the Q4_K_M file size, add 1-1.5GB for overhead, and compare against your available VRAM. If it fits with room to spare, you will have a good experience. If it barely fits, expect limited context windows and occasional slowdowns.


Building a new machine around a specific model? Start with the hardware requirements guide to size your GPU, RAM, and storage correctly.

🎯
AI Learning Path

Ollama’s running. Here’s what to build with it.

Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.

Or own it for life — Lifetime $149 $599, pay once
Once your hardware is sorted

Stop piecing Ollama together from blog posts

Ollama Mastery is 15 chapters end to end — install, model choice, Modelfiles, GPU offload, the API, and the 20 errors that actually happen. Plus 24 more courses.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsOllamaRAMVRAMModelsQuantizationHardware

Local AI Master Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on Local AI Hardware
See the full AI Hardware Guide 2026 guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

📅 Published: April 11, 2026🔄 Last Updated: September 15, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor

Get Model Updates Weekly

New Ollama models drop every week. Get VRAM requirements, sizing math, and recommendations before everyone else.

Was this helpful?

📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Ollama’s running. Here’s what to build with it.

Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators