Models

Best LLM for 64GB RAM: What a CPU-Only Box Runs

October 4, 2026
11 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Got the hardware sorted? Now build on it. You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Start free
Or own it for life — Lifetime $149, pay once

Short answer: with 64GB of system RAM, pull qwen3-next:80b — a 50.35GB Q4_K_M download that activates about 3B of its 79.7B parameters per token, so a CPU reads roughly 2GB of weights per token instead of 50. That is an 80-billion-parameter model running at the speed of a 3B one. A dense llama3.3:70b is smaller on disk at 42.52GB and lands near 1.5 tokens per second on dual-channel DDR5 — around twenty times slower.

So yes, 64GB unlocks a genuinely bigger class of model. But it unlocks it in one specific shape: mixture-of-experts, not dense. Buy the capacity, spend it on sparsity. Everything below is the arithmetic behind that sentence.

Every download size here was read from the Ollama registry manifests on 18 August 2026 — exact bytes, not rounded marketing figures. Every tokens-per-second number is a calculated ceiling from the bandwidth formula, labelled as such, because we would rather show you the working than quote a benchmark that was not run on your memory.

If you are on 32GB, the 32GB tier page is the one you want — the answers there are different and the model list does not overlap with this one.


What 64GB Actually Buys

Not speed. Capacity — and a specific loophole for spending it.

Token generation on a CPU is a memory-streaming problem. To emit one token the machine has to read every weight the model needs for that token out of RAM. So tokens per second tracks memory bandwidth, and bandwidth does not improve when you add DIMM capacity. A 64GB box and a 32GB box with the same memory kit run the same model at the same speed.

What changes is what will fit. And in 2026 that matters more than it used to, because mixture-of-experts models decoupled "how big is it" from "how much of it do I read":

  • A dense model reads all of itself per token. Bigger means proportionally slower, always.
  • A mixture-of-experts model routes each token through a small subset of its experts. It must live in RAM in full — the router can pick any expert next — but it only reads a fraction of itself per token.

That is why the headline pick here is 50GB and fast while a 42GB model is slow. MoE converts spare capacity into quality without spending speed, and 64GB is the first tier where you have enough spare capacity to do it properly.

One accounting note that trips people up: a "64GB" memory kit is 64 GiB, which is about 68.7GB in the decimal units Ollama uses when it reports sizes. It sounds pedantic until you are looking at a 65.37GB model wondering why it will not run. See what does not fit.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

The Models

Sizes are exact bytes from the Ollama registry manifest; parameter counts and quantisation come from each model's config blob. The tokens-per-second columns are calculated ceilings, explained in the next section.

#Pull commandDownloadParams / quantRead per tokenDDR4-3200DDR5-6000
1ollama pull qwen3-next:80b50.35 GB79.7B MoE, Q4_K_M~2.0 GB~17 tok/s~31 tok/s
2ollama pull qwen3-coder:30b18.56 GB30.5B MoE, Q4_K_M~1.9 GB~17 tok/s~33 tok/s
3ollama pull mixtral:8x7b26.44 GB46.7B MoE, Q4_0~7.2 GB~5 tok/s~9 tok/s
4ollama pull gemma3:27b17.40 GB27.4B dense, Q4_K_M17.4 GB~1.9 tok/s~3.6 tok/s
5ollama pull qwen3:32b20.20 GB32.8B dense, Q4_K_M20.2 GB~1.6 tok/s~3.1 tok/s
6ollama pull llama3.3:70b42.52 GB70.6B dense, Q4_K_M42.5 GB~0.8 tok/s~1.5 tok/s
7ollama pull qwen2.5:72b47.42 GB72.7B dense, Q4_K_M47.4 GB~0.7 tok/s~1.3 tok/s

Manifest sizes and config metadata read from registry.ollama.ai on 18 August 2026. Both bandwidth columns assume two DIMMs at 65% of theoretical peak — see The Speed Math.

How to read that ordering:

  • Pick 1 for the best model your box can actually converse with. Qwen3-Next-80B-A3B is the largest thing on the list and the second fastest, and that ordering is the whole point of this page.
  • Pick 2 if you code. qwen3-coder:30b is the same A3B sparsity in a third of the footprint, which means it leaves ~50GB free for everything else you are doing. At 64GB you can genuinely keep this loaded all day.
  • Pick 3 if you want the classic MoE. Mixtral 8x7B routes 2 of 8 experts per token, and its shared attention means about 12.7B active parameters — far more than the A3B models, hence the drop to single digits. It is also an older model; it is on the list because it is the sparsity mid-point and it shows the pattern is not magic.
  • Picks 4-7 are dense, and dense is the slow lane. They earn their place only when you specifically want what they are: gemma3:27b for multilingual and multimodal work, qwen3:32b as a strong general dense model, llama3.3:70b and qwen2.5:72b as the smartest weights that fit — at batch-job speeds.

A caveat on pick 1 that we will not bury. Qwen3-Next uses a hybrid attention design — only 12 of its 48 layers are full attention, the rest use a linear/gated-delta mechanism. That is excellent for memory (see context) but it is more compute-heavy per token than a plain MoE transformer, so real throughput will fall further below the bandwidth ceiling than it does for the smaller Qwen3 MoE models. It is also a newer architecture in the local runtimes. Pull it, run --verbose, and believe your own number over ours.


The Speed Math

tokens/second ≈ effective memory bandwidth ÷ bytes read per token. Every figure in the table above is that one line, and it is why your DIMMs matter more than your CPU.

Theoretical peak bandwidth is arithmetic: transfers per second × 8 bytes per channel × channel count.

ConfigurationTheoretical peak~65% effective
DDR4-3200, 2 channels51.2 GB/s~33 GB/s
DDR5-5600, 2 channels89.6 GB/s~58 GB/s
DDR5-6000, 2 channels96.0 GB/s~62 GB/s
DDR4-3200, 4 channels (HEDT)102.4 GB/s~67 GB/s
DDR5-5200, 4 channels166.4 GB/s~108 GB/s
DDR5-4800, 8 channels (workstation/server)307.2 GB/s~200 GB/s

The 65% is an assumption, not a measurement. Real streaming workloads land somewhere around 60-75% of peak depending on the memory controller, the timings and how many DIMMs are populated. We picked one number and applied it consistently so the models compare fairly against each other. The ratios are the durable part.

Three things specific to the 64GB tier, and they matter more here than at 32GB because 64GB buyers are the ones most likely to have an unusual memory layout:

Four sticks are usually slower than two on DDR5. A 64GB kit can be 2×32GB or 4×16GB. On DDR5, populating all four slots frequently forces the controller down to a lower speed grade — a kit rated 6000 may end up training at 4400 or lower. That is a real, silent ~25% throughput loss on every model in the table. If you have a choice at purchase time, buy two 32GB sticks.

Check XMP or EXPO is on. A DDR5-6000 kit running at its JEDEC default of 4800 gives you 76.8 GB/s instead of 96 — you paid for the faster kit and are not using it. This is one BIOS toggle.

Quad and eight-channel platforms are the only real speed upgrade. A Threadripper or EPYC box with eight populated channels has more than three times the bandwidth of a good consumer desktop, and that is the one change that moves a dense 70B out of the unusable range. No consumer CPU upgrade does anything comparable, because the cores are already idle waiting on memory.

Prompt processing behaves differently. Reading your prompt is compute-bound and parallel; generating is memory-bound and serial. A CPU chews through a long pasted document far faster than it writes a reply, so expect an up-front pause and then a steady slow crawl. Ollama reports the two rates separately, which is how you should read your own results. Our CPU-only inference guide has the third-party measurements we use to sanity-check this model.


The 70B Question

A dense 70B fits in 64GB and generates at roughly 0.8-1.5 tokens per second on a normal desktop. That is real, it is usable for some things, and it is not a chat model.

This is the claim most pages at this tier get wrong by omission. "64GB runs a 70B" is true. What it means in practice:

TaskTokensAt 1.2 tok/s
A one-line answer~40~35 seconds
A normal paragraph reply~150~2 minutes
A detailed explanation~500~7 minutes
A long document draft~2,000~28 minutes

Compare the same tasks on qwen3-next:80b at a ceiling near 31 tok/s: the 500-token answer takes about 16 seconds. Same box, same RAM, a bigger model, and the difference is entirely which parameters get read.

So when is a dense 70B still the right call? When quality per token matters more than tokens per hour and you are not sitting there watching:

  • Overnight batch work — summarising a folder, generating a dataset, running an evaluation.
  • A single hard question where you will go and do something else.
  • Any workflow where you were going to queue requests anyway.

When it is the wrong call: interactive chat, coding assistance, anything with a person waiting. For those, pick 1 or pick 2 from the table, and accept that the 70B's extra reasoning is not worth thirty minutes of your afternoon.

If a dense 70B is genuinely your requirement, the fix is bandwidth, not capacity. An eight-channel workstation takes it to roughly 4-5 tok/s by the same formula — still slow, but a different category of slow. And if you have a graphics card at all, the calculus changes completely: start at the VRAM calculator or the 32GB VRAM model picks instead, because GPU memory is five to fifteen times faster than anything in the table above.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

What Does Not Fit

Three models people specifically arrive at this tier hoping for, and the honest verdict on each.

ModelDownloadVerdict at 64GB
gpt-oss:120b65.37 GBNo. Leaves ~3GB for OS, runtime and KV cache
llama4:scout67.44 GBNo. Same problem, slightly worse
command-r-plus:104b59.22 GBTechnically yes, with a short context and nothing else running
qwen3:235b-a22b142.15 GBNot close
deepseek-v2.5:236b132.91 GBNot close

The gpt-oss:120b case deserves the detail because it is the one people ask about. It is a 116.8B MXFP4 model with 128 experts and 4 active per token — architecturally, exactly the kind of model this tier wants. But 65.37GB of weights against 68.7GB of addressable memory leaves no room for the operating system, the runtime's compute buffers, or a single token of KV cache. It will load via memory mapping and then thrash against your SSD, and once weights are coming off disk your effective bandwidth is an order of magnitude below DDR4.

This is the single clearest argument for 96GB or 128GB rather than 64GB, and it is worth weighing before you buy. 32 vs 64 vs 128GB for local AI works through that purchase decision properly; this page assumes you already have the 64GB.


Context Is Nearly Free (On One Model)

KV cache is the second half of your memory budget, and it varies by a factor of thirteen between the models in the table.

The cache grows linearly with context length and you can compute it exactly from the architecture:

KV bytes per token = 2 (K and V) × attention layers × kv_heads × head_dim × bytes_per_element

Run that for the two ends of the table, at fp16:

ModelAttention layersKV headsHead dimPer token32K context128K context
qwen3-next:80b12 of 48 (hybrid)225624 KiB~0.8 GB~3.2 GB
qwen3:32b648128256 KiB~8.6 GB~34 GB
llama3.3:70b808128320 KiB~10.7 GB~43 GB

Computed from the published config.json for each model. Qwen3-Next uses full attention on only every fourth layer; the remaining layers hold a fixed-size recurrent state that does not grow with context.

Read the right-hand column carefully, because it changes what these models are for:

  • qwen3-next:80b at 128K context needs 50.35 + 3.2 ≈ 53.6GB. That comfortably fits in 64GB with the OS. A 128K-context local model on a CPU box is a genuinely new capability at this tier.
  • llama3.3:70b at 128K context needs 42.52 + 43 ≈ 85.5GB. Impossible. Even 32K puts you at ~53GB before compute buffers, which is workable but leaves little slack.

So the honest framing of the original question — bigger model or bigger context? — is that at 64GB you can have both, but only from a model whose attention was designed for it.

Two levers if you are tight:

# Halve the KV cache (default is f16)
OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve
# Pin the context instead of letting it be chosen for you
OLLAMA_CONTEXT_LENGTH=32768 ollama serve

Ollama's OLLAMA_CONTEXT_LENGTH defaults to 0, which its own environment documentation describes as "4k/32k/256k based on VRAM" — on a CPU-only box you want that decision to be explicit. OLLAMA_KV_CACHE_TYPE defaults to f16. Also leave OLLAMA_NUM_PARALLEL at its default of 1: each parallel slot gets its own KV allocation, and on a bandwidth-bound machine two concurrent requests run at half speed each while doubling your memory use. For the full cross-model lookup, the Ollama model RAM and VRAM table and RAM requirements for local AI cover the reverse direction.


Measure Your Own Box

Do not trust our ceilings. Two commands give you the real number for your memory in about a minute, and at this tier the spread between machines is enormous.

# Real tokens per second on your hardware
ollama run qwen3-next:80b --verbose

Ask something that needs a few hundred tokens of answer. The summary afterwards separates prompt eval rate from eval rate. The second is the one that decides whether you keep using the model.

# What is resident, and confirm it is all on CPU
ollama ps

PROCESSOR should read 100% CPU. SIZE is your true footprint including the KV cache — compare it against your free memory, not against the download size.

If you want to compare quantisations, thread counts or memory configurations properly, llama-bench from llama.cpp runs a fixed workload and reports prompt and generation rates separately, which makes the comparison fair. That is the tool to use if you are deciding whether four sticks or two is faster on your specific board — a question no published table can answer for you.


The Usability Line

Below 5 tokens per second, a local model is a batch tool. Above 10, it is something you will open again tomorrow. At 64GB you can land on either side of that line depending purely on which model you pull.

  • Under 1 tok/s — the dense 70B/72B tier on a consumer desktop. Overnight work only.
  • 1-3 tok/s — dense 27B-32B on DDR5, or a 70B on a quad-channel box. You will watch the cursor.
  • 3-5 tok/s — dense 27B on fast DDR5, or a 70B on eight channels. Tolerable for a single question.
  • 5-10 tok/s — Mixtral-class sparsity. Roughly reading pace.
  • 15-35 tok/s — the A3B mixture-of-experts models. Genuinely comfortable, and the reason to have 64GB.

The bottom line for this tier: 64GB is a real local-AI machine, and the thing it unlocks is an 80B model at 30-billion-model speed — not a 70B dense model at usable speed. If you buy 64GB expecting the second, you will conclude that CPU inference does not work, when the actual finding was that you picked the model that reads twenty times more memory per token.

And if you are still choosing hardware rather than models, our hardware hub and the mini PC roundup cover the boxes where unified memory changes this arithmetic again.


Sources

  • Ollama registry (registry.ollama.ai) — every download size, parameter count and quantisation type on this page was read from the model manifests and config blobs on 18 August 2026. Exact bytes; sizes shown in decimal GB, matching how Ollama reports them.
  • Qwen3-Next-80B-A3B config.json (Hugging Face, Qwen/Qwen3-Next-80B-A3B-Instruct) — 48 layers, 512 experts with 10 active plus 1 shared, full_attention_interval 4, 2 KV heads, head dim 256. Source of the active-parameter and KV-cache figures.
  • openai/gpt-oss-120b config.json — 36 layers, 128 local experts, 4 active per token.
  • Llama 3.3 70B config.json — 80 layers, 8 KV heads, head dim 128. Mixtral-8x7B config.json — 32 layers, 8 experts, 2 active per token. Qwen3-32B config.json — 64 layers, 8 KV heads.
  • ollama/ollama source, envconfig/config.go — OLLAMA_CONTEXT_LENGTH default 0 ("4k/32k/256k based on VRAM"), OLLAMA_KV_CACHE_TYPE default f16, OLLAMA_NUM_PARALLEL default 1.
  • Third-party CPU measurements — the published bandwidth-scaling and CPU-only throughput figures cited and linked in our CPU-only inference guide, used here only to sanity-check the formula rather than as measurements of these models.

FAQ

🎯
AI Learning Path

Got the hardware sorted? Now build on it.

You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Once your hardware is sorted

Decide before you spend a thousand pounds

The AI Hardware course sizes your build properly — VRAM ladder, real bottlenecks, budget builds — and Pick the Right Model tells you what to run on it.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
Tags64GB RAMCPU InferenceNo GPUOllamaQwen3-NextMoEMemory Bandwidth

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on Local AI Hardware
See the full AI Hardware Guide 2026 guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

What is the best LLM to run with 64GB of RAM and no GPU?

qwen3-next:80b. It is a 50.35GB Q4_K_M download (79.7B total parameters, read from the Ollama registry manifest on 18 August 2026) and it is a mixture-of-experts model that activates roughly 3B parameters per token. On a CPU, where speed is set by how many bytes you read per token, that makes an 80-billion-parameter model roughly twenty times faster than a dense 70B of similar size on disk. If you want a dense model for its own sake, qwen3:32b is 20.2GB and lands around 3 tokens per second on DDR5.

Can 64GB of RAM run a 70B model?

Yes, and you will not enjoy it. llama3.3:70b at Q4_K_M is a 42.52GB download, so it loads with room to spare. But a dense model reads essentially all of itself for every token it emits, so on dual-channel DDR5-6000 the arithmetic gives a ceiling near 1.5 tokens per second, and on DDR4-3200 nearer 0.8. A 500-token answer at 1.2 tok/s takes about seven minutes. It works, it is genuinely the smartest thing your box can hold, and it is a batch tool rather than a chat tool.

Does 64GB of RAM let me run gpt-oss:120b?

No, not comfortably. The Ollama manifest puts gpt-oss:120b at 65.37GB. A "64GB" kit is 64 GiB, which is about 68.7GB in the decimal units Ollama reports, so the weights alone would leave roughly 3GB for Windows or Linux, the runtime, compute buffers and the KV cache. It will page to disk, and once you are reading weights off an SSD your effective bandwidth collapses. This model is the reason 96GB and 128GB kits exist. The same applies to llama4:scout at 67.44GB.

Is 64GB better than 32GB for local LLMs on a CPU?

It buys you a better model at the same speed, not a faster machine. Both tiers are limited by memory bandwidth, which does not change when you add capacity. What 64GB adds is the room to hold a much larger mixture-of-experts model whose active-parameter count — and therefore its speed — is similar to what a 32GB box was already running. Going from qwen3:30b-a3b to qwen3-next:80b is a real quality jump at roughly unchanged tokens per second. Going from a 14B dense to a 70B dense is a quality jump at a quarter of the speed.

Does a Threadripper or EPYC box change the answer?

More than any consumer upgrade does, because it changes the one variable that matters. Dual-channel DDR5-6000 peaks at 96 GB/s; an eight-channel DDR5-4800 platform peaks at 307.2 GB/s, more than three times as much. That moves a dense 70B from around 1.5 tok/s to something in the mid single digits, which is the difference between a batch job and a slow conversation. If you have a quad-channel or eight-channel machine, run the measurement commands at the end of this page rather than trusting any table — unusual memory topologies are exactly where published figures stop applying.

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks for local models — skip the setup and start pulling. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: October 4, 2026🔄 Last Updated: October 4, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
Free Tools & Calculators