Best LLM for 64GB RAM: What a CPU-Only Box Runs
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Got the hardware sorted? Now build on it. You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.
Short answer: with 64GB of system RAM, pull qwen3-next:80b — a 50.35GB Q4_K_M download that activates about 3B of its 79.7B parameters per token, so a CPU reads roughly 2GB of weights per token instead of 50. That is an 80-billion-parameter model running at the speed of a 3B one. A dense llama3.3:70b is smaller on disk at 42.52GB and lands near 1.5 tokens per second on dual-channel DDR5 — around twenty times slower.
So yes, 64GB unlocks a genuinely bigger class of model. But it unlocks it in one specific shape: mixture-of-experts, not dense. Buy the capacity, spend it on sparsity. Everything below is the arithmetic behind that sentence.
Every download size here was read from the Ollama registry manifests on 18 August 2026 — exact bytes, not rounded marketing figures. Every tokens-per-second number is a calculated ceiling from the bandwidth formula, labelled as such, because we would rather show you the working than quote a benchmark that was not run on your memory.
If you are on 32GB, the 32GB tier page is the one you want — the answers there are different and the model list does not overlap with this one.
What 64GB Actually Buys
Not speed. Capacity — and a specific loophole for spending it.
Token generation on a CPU is a memory-streaming problem. To emit one token the machine has to read every weight the model needs for that token out of RAM. So tokens per second tracks memory bandwidth, and bandwidth does not improve when you add DIMM capacity. A 64GB box and a 32GB box with the same memory kit run the same model at the same speed.
What changes is what will fit. And in 2026 that matters more than it used to, because mixture-of-experts models decoupled "how big is it" from "how much of it do I read":
- A dense model reads all of itself per token. Bigger means proportionally slower, always.
- A mixture-of-experts model routes each token through a small subset of its experts. It must live in RAM in full — the router can pick any expert next — but it only reads a fraction of itself per token.
That is why the headline pick here is 50GB and fast while a 42GB model is slow. MoE converts spare capacity into quality without spending speed, and 64GB is the first tier where you have enough spare capacity to do it properly.
One accounting note that trips people up: a "64GB" memory kit is 64 GiB, which is about 68.7GB in the decimal units Ollama uses when it reports sizes. It sounds pedantic until you are looking at a 65.37GB model wondering why it will not run. See what does not fit.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
The Models
Sizes are exact bytes from the Ollama registry manifest; parameter counts and quantisation come from each model's config blob. The tokens-per-second columns are calculated ceilings, explained in the next section.
| # | Pull command | Download | Params / quant | Read per token | DDR4-3200 | DDR5-6000 |
|---|---|---|---|---|---|---|
| 1 | ollama pull qwen3-next:80b | 50.35 GB | 79.7B MoE, Q4_K_M | ~2.0 GB | ~17 tok/s | ~31 tok/s |
| 2 | ollama pull qwen3-coder:30b | 18.56 GB | 30.5B MoE, Q4_K_M | ~1.9 GB | ~17 tok/s | ~33 tok/s |
| 3 | ollama pull mixtral:8x7b | 26.44 GB | 46.7B MoE, Q4_0 | ~7.2 GB | ~5 tok/s | ~9 tok/s |
| 4 | ollama pull gemma3:27b | 17.40 GB | 27.4B dense, Q4_K_M | 17.4 GB | ~1.9 tok/s | ~3.6 tok/s |
| 5 | ollama pull qwen3:32b | 20.20 GB | 32.8B dense, Q4_K_M | 20.2 GB | ~1.6 tok/s | ~3.1 tok/s |
| 6 | ollama pull llama3.3:70b | 42.52 GB | 70.6B dense, Q4_K_M | 42.5 GB | ~0.8 tok/s | ~1.5 tok/s |
| 7 | ollama pull qwen2.5:72b | 47.42 GB | 72.7B dense, Q4_K_M | 47.4 GB | ~0.7 tok/s | ~1.3 tok/s |
Manifest sizes and config metadata read from registry.ollama.ai on 18 August 2026. Both bandwidth columns assume two DIMMs at 65% of theoretical peak — see The Speed Math.
How to read that ordering:
- Pick 1 for the best model your box can actually converse with. Qwen3-Next-80B-A3B is the largest thing on the list and the second fastest, and that ordering is the whole point of this page.
- Pick 2 if you code.
qwen3-coder:30bis the same A3B sparsity in a third of the footprint, which means it leaves ~50GB free for everything else you are doing. At 64GB you can genuinely keep this loaded all day. - Pick 3 if you want the classic MoE. Mixtral 8x7B routes 2 of 8 experts per token, and its shared attention means about 12.7B active parameters — far more than the A3B models, hence the drop to single digits. It is also an older model; it is on the list because it is the sparsity mid-point and it shows the pattern is not magic.
- Picks 4-7 are dense, and dense is the slow lane. They earn their place only when you specifically want what they are:
gemma3:27bfor multilingual and multimodal work,qwen3:32bas a strong general dense model,llama3.3:70bandqwen2.5:72bas the smartest weights that fit — at batch-job speeds.
A caveat on pick 1 that we will not bury. Qwen3-Next uses a hybrid attention design — only 12 of its 48 layers are full attention, the rest use a linear/gated-delta mechanism. That is excellent for memory (see context) but it is more compute-heavy per token than a plain MoE transformer, so real throughput will fall further below the bandwidth ceiling than it does for the smaller Qwen3 MoE models. It is also a newer architecture in the local runtimes. Pull it, run --verbose, and believe your own number over ours.
The Speed Math
tokens/second ≈ effective memory bandwidth ÷ bytes read per token. Every figure in the table above is that one line, and it is why your DIMMs matter more than your CPU.
Theoretical peak bandwidth is arithmetic: transfers per second × 8 bytes per channel × channel count.
| Configuration | Theoretical peak | ~65% effective |
|---|---|---|
| DDR4-3200, 2 channels | 51.2 GB/s | ~33 GB/s |
| DDR5-5600, 2 channels | 89.6 GB/s | ~58 GB/s |
| DDR5-6000, 2 channels | 96.0 GB/s | ~62 GB/s |
| DDR4-3200, 4 channels (HEDT) | 102.4 GB/s | ~67 GB/s |
| DDR5-5200, 4 channels | 166.4 GB/s | ~108 GB/s |
| DDR5-4800, 8 channels (workstation/server) | 307.2 GB/s | ~200 GB/s |
The 65% is an assumption, not a measurement. Real streaming workloads land somewhere around 60-75% of peak depending on the memory controller, the timings and how many DIMMs are populated. We picked one number and applied it consistently so the models compare fairly against each other. The ratios are the durable part.
Three things specific to the 64GB tier, and they matter more here than at 32GB because 64GB buyers are the ones most likely to have an unusual memory layout:
Four sticks are usually slower than two on DDR5. A 64GB kit can be 2×32GB or 4×16GB. On DDR5, populating all four slots frequently forces the controller down to a lower speed grade — a kit rated 6000 may end up training at 4400 or lower. That is a real, silent ~25% throughput loss on every model in the table. If you have a choice at purchase time, buy two 32GB sticks.
Check XMP or EXPO is on. A DDR5-6000 kit running at its JEDEC default of 4800 gives you 76.8 GB/s instead of 96 — you paid for the faster kit and are not using it. This is one BIOS toggle.
Quad and eight-channel platforms are the only real speed upgrade. A Threadripper or EPYC box with eight populated channels has more than three times the bandwidth of a good consumer desktop, and that is the one change that moves a dense 70B out of the unusable range. No consumer CPU upgrade does anything comparable, because the cores are already idle waiting on memory.
Prompt processing behaves differently. Reading your prompt is compute-bound and parallel; generating is memory-bound and serial. A CPU chews through a long pasted document far faster than it writes a reply, so expect an up-front pause and then a steady slow crawl. Ollama reports the two rates separately, which is how you should read your own results. Our CPU-only inference guide has the third-party measurements we use to sanity-check this model.
The 70B Question
A dense 70B fits in 64GB and generates at roughly 0.8-1.5 tokens per second on a normal desktop. That is real, it is usable for some things, and it is not a chat model.
This is the claim most pages at this tier get wrong by omission. "64GB runs a 70B" is true. What it means in practice:
| Task | Tokens | At 1.2 tok/s |
|---|---|---|
| A one-line answer | ~40 | ~35 seconds |
| A normal paragraph reply | ~150 | ~2 minutes |
| A detailed explanation | ~500 | ~7 minutes |
| A long document draft | ~2,000 | ~28 minutes |
Compare the same tasks on qwen3-next:80b at a ceiling near 31 tok/s: the 500-token answer takes about 16 seconds. Same box, same RAM, a bigger model, and the difference is entirely which parameters get read.
So when is a dense 70B still the right call? When quality per token matters more than tokens per hour and you are not sitting there watching:
- Overnight batch work — summarising a folder, generating a dataset, running an evaluation.
- A single hard question where you will go and do something else.
- Any workflow where you were going to queue requests anyway.
When it is the wrong call: interactive chat, coding assistance, anything with a person waiting. For those, pick 1 or pick 2 from the table, and accept that the 70B's extra reasoning is not worth thirty minutes of your afternoon.
If a dense 70B is genuinely your requirement, the fix is bandwidth, not capacity. An eight-channel workstation takes it to roughly 4-5 tok/s by the same formula — still slow, but a different category of slow. And if you have a graphics card at all, the calculus changes completely: start at the VRAM calculator or the 32GB VRAM model picks instead, because GPU memory is five to fifteen times faster than anything in the table above.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
What Does Not Fit
Three models people specifically arrive at this tier hoping for, and the honest verdict on each.
| Model | Download | Verdict at 64GB |
|---|---|---|
gpt-oss:120b | 65.37 GB | No. Leaves ~3GB for OS, runtime and KV cache |
llama4:scout | 67.44 GB | No. Same problem, slightly worse |
command-r-plus:104b | 59.22 GB | Technically yes, with a short context and nothing else running |
qwen3:235b-a22b | 142.15 GB | Not close |
deepseek-v2.5:236b | 132.91 GB | Not close |
The gpt-oss:120b case deserves the detail because it is the one people ask about. It is a 116.8B MXFP4 model with 128 experts and 4 active per token — architecturally, exactly the kind of model this tier wants. But 65.37GB of weights against 68.7GB of addressable memory leaves no room for the operating system, the runtime's compute buffers, or a single token of KV cache. It will load via memory mapping and then thrash against your SSD, and once weights are coming off disk your effective bandwidth is an order of magnitude below DDR4.
This is the single clearest argument for 96GB or 128GB rather than 64GB, and it is worth weighing before you buy. 32 vs 64 vs 128GB for local AI works through that purchase decision properly; this page assumes you already have the 64GB.
Context Is Nearly Free (On One Model)
KV cache is the second half of your memory budget, and it varies by a factor of thirteen between the models in the table.
The cache grows linearly with context length and you can compute it exactly from the architecture:
KV bytes per token = 2 (K and V) × attention layers × kv_heads × head_dim × bytes_per_element
Run that for the two ends of the table, at fp16:
| Model | Attention layers | KV heads | Head dim | Per token | 32K context | 128K context |
|---|---|---|---|---|---|---|
qwen3-next:80b | 12 of 48 (hybrid) | 2 | 256 | 24 KiB | ~0.8 GB | ~3.2 GB |
qwen3:32b | 64 | 8 | 128 | 256 KiB | ~8.6 GB | ~34 GB |
llama3.3:70b | 80 | 8 | 128 | 320 KiB | ~10.7 GB | ~43 GB |
Computed from the published config.json for each model. Qwen3-Next uses full attention on only every fourth layer; the remaining layers hold a fixed-size recurrent state that does not grow with context.
Read the right-hand column carefully, because it changes what these models are for:
qwen3-next:80bat 128K context needs 50.35 + 3.2 ≈ 53.6GB. That comfortably fits in 64GB with the OS. A 128K-context local model on a CPU box is a genuinely new capability at this tier.llama3.3:70bat 128K context needs 42.52 + 43 ≈ 85.5GB. Impossible. Even 32K puts you at ~53GB before compute buffers, which is workable but leaves little slack.
So the honest framing of the original question — bigger model or bigger context? — is that at 64GB you can have both, but only from a model whose attention was designed for it.
Two levers if you are tight:
# Halve the KV cache (default is f16)
OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve
# Pin the context instead of letting it be chosen for you
OLLAMA_CONTEXT_LENGTH=32768 ollama serve
Ollama's OLLAMA_CONTEXT_LENGTH defaults to 0, which its own environment documentation describes as "4k/32k/256k based on VRAM" — on a CPU-only box you want that decision to be explicit. OLLAMA_KV_CACHE_TYPE defaults to f16. Also leave OLLAMA_NUM_PARALLEL at its default of 1: each parallel slot gets its own KV allocation, and on a bandwidth-bound machine two concurrent requests run at half speed each while doubling your memory use. For the full cross-model lookup, the Ollama model RAM and VRAM table and RAM requirements for local AI cover the reverse direction.
Measure Your Own Box
Do not trust our ceilings. Two commands give you the real number for your memory in about a minute, and at this tier the spread between machines is enormous.
# Real tokens per second on your hardware
ollama run qwen3-next:80b --verbose
Ask something that needs a few hundred tokens of answer. The summary afterwards separates prompt eval rate from eval rate. The second is the one that decides whether you keep using the model.
# What is resident, and confirm it is all on CPU
ollama ps
PROCESSOR should read 100% CPU. SIZE is your true footprint including the KV cache — compare it against your free memory, not against the download size.
If you want to compare quantisations, thread counts or memory configurations properly, llama-bench from llama.cpp runs a fixed workload and reports prompt and generation rates separately, which makes the comparison fair. That is the tool to use if you are deciding whether four sticks or two is faster on your specific board — a question no published table can answer for you.
The Usability Line
Below 5 tokens per second, a local model is a batch tool. Above 10, it is something you will open again tomorrow. At 64GB you can land on either side of that line depending purely on which model you pull.
- Under 1 tok/s — the dense 70B/72B tier on a consumer desktop. Overnight work only.
- 1-3 tok/s — dense 27B-32B on DDR5, or a 70B on a quad-channel box. You will watch the cursor.
- 3-5 tok/s — dense 27B on fast DDR5, or a 70B on eight channels. Tolerable for a single question.
- 5-10 tok/s — Mixtral-class sparsity. Roughly reading pace.
- 15-35 tok/s — the A3B mixture-of-experts models. Genuinely comfortable, and the reason to have 64GB.
The bottom line for this tier: 64GB is a real local-AI machine, and the thing it unlocks is an 80B model at 30-billion-model speed — not a 70B dense model at usable speed. If you buy 64GB expecting the second, you will conclude that CPU inference does not work, when the actual finding was that you picked the model that reads twenty times more memory per token.
And if you are still choosing hardware rather than models, our hardware hub and the mini PC roundup cover the boxes where unified memory changes this arithmetic again.
Sources
- Ollama registry (
registry.ollama.ai) — every download size, parameter count and quantisation type on this page was read from the model manifests and config blobs on 18 August 2026. Exact bytes; sizes shown in decimal GB, matching how Ollama reports them. - Qwen3-Next-80B-A3B
config.json(Hugging Face,Qwen/Qwen3-Next-80B-A3B-Instruct) — 48 layers, 512 experts with 10 active plus 1 shared,full_attention_interval4, 2 KV heads, head dim 256. Source of the active-parameter and KV-cache figures. openai/gpt-oss-120bconfig.json— 36 layers, 128 local experts, 4 active per token.- Llama 3.3 70B
config.json— 80 layers, 8 KV heads, head dim 128. Mixtral-8x7Bconfig.json— 32 layers, 8 experts, 2 active per token. Qwen3-32Bconfig.json— 64 layers, 8 KV heads. - ollama/ollama source,
envconfig/config.go—OLLAMA_CONTEXT_LENGTHdefault 0 ("4k/32k/256k based on VRAM"),OLLAMA_KV_CACHE_TYPEdefault f16,OLLAMA_NUM_PARALLELdefault 1. - Third-party CPU measurements — the published bandwidth-scaling and CPU-only throughput figures cited and linked in our CPU-only inference guide, used here only to sanity-check the formula rather than as measurements of these models.
FAQ
Got the hardware sorted? Now build on it.
You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.
Decide before you spend a thousand pounds
The AI Hardware course sizes your build properly — VRAM ladder, real bottlenecks, budget builds — and Pick the Right Model tells you what to run on it.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARLocal AI Hardware Requirements (2026): Complete Guide
- AI Hardware Requirements: CPU, GPU and RAM for Beginners
- AI RAM Requirements 2026: How Much for 7B, 13B, 70B Models?
- AI Server Build Under $1,500: Parts List and What Fits
- AMD GPU Not Supported by ROCm? HSA_OVERRIDE Values
- AMD MI50 32GB for Local LLMs: The Used VRAM King, Honestly
- AMD Ryzen AI Max+ 395 (Strix Halo) for Local AI 2026
- Apple M4 for Local AI: Mac Studio + MacBook Guide (2026)
- Benchmark Your Local AI Setup: tok/s, TTFT, VRAM
- Best GPU for AI Video Generation: By VRAM Tier (2026)
Comments (0)
No comments yet. Be the first to share your thoughts!