★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once

Llama 3 70B:
Technical Analysis & Setup

Updated: September 29, 2026

Llama 3 70B is the larger of the two models Meta released on 18 April 2024. It has 70.6 billion parameters and an 8K context window, and the default Ollama build, llama3:70b, is a 40 GB download. That puts it on 48 GB of VRAM, or on one 32 GB card at 2-bit quantisation. It has been replaced by Llama 3.1 70B and Llama 3.3 70B, which are the same size with a 128K context. Everything on this page is copied from Meta's model card and licence, and from the file sizes published on Ollama and Hugging Face; this site has not benchmarked the model.

70.6B
Parameters
8K
Context window
40 GB
Ollama llama3:70b download
82.0
MMLU, instruct model (Meta)

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

Llama 3 70B hardware requirements

Llama 3 70B needs about 46 GB of memory at Q4_K_M with its full 8K context, so the standard setup is 48 GB of VRAM: two 24 GB cards or one 48 GB card. The 26 GB Q2_K build fits a single 32 GB card. A single 24 GB card holds only the 19 GB IQ2_XXS build, or runs a larger build with part of the model in system RAM. The arithmetic is parameters × bits per weight ÷ 8: 70.55B × 4.82 bits ÷ 8 = 42.5 GB at Q4_K_M.

File sizes are from bartowski/Meta-Llama-3-70B-Instruct-GGUF and Ollama's tag list. The context cache is arithmetic from the model's config.json: 80 layers × 8 key-value heads × 128 dimensions, stored twice at 2 bytes, is 328 KB per token, or 2.7 GB for 8,192 tokens. “Memory needed” is file + that cache + a 1 GB allowance for buffers.

QuantisationFileOllama tag (listed size)Memory needed (8K context)Fits
IQ2_XXS19.10 GBnot on Ollama~22.8 GBOne 24 GB card (RTX 3090, RTX 4090), with heavy quality loss
Q2_K26.38 GB70b-instruct-q2_K (26GB)~30.1 GBOne 32 GB card (RTX 5090)
Q3_K_M34.27 GB70b-instruct-q3_K_M (34GB)~38.0 GB48 GB: two 24 GB cards or one 48 GB card
q4_040 GB (Ollama)llama3:70b (40GB)~43.7 GB48 GB: two 24 GB cards or one 48 GB card
Q4_K_M42.52 GB70b-instruct-q4_K_M (43GB)~46.2 GB48 GB, with under 2 GB to spare
Q5_K_M49.95 GB70b-instruct-q5_K_M (50GB)~53.7 GB64 GB: two 32 GB cards or three 24 GB cards
Q6_K57.89 GB70b-instruct-q6_K (58GB)~61.6 GB64 GB, tight; 72 GB across three 24 GB cards
Q8_075 GB (Ollama)70b-instruct-q8_0 (75GB)~78.7 GB96 GB
fp16141.12 GB70b-instruct-fp16 (141GB)~144.8 GBTwo 80 GB data-centre cards
  • One 24 GB card: the 40 GB default does not fit. Ollama will split the model between the card and system RAM, which works but is slower than running fully on the GPU; you need about 48 GB of combined memory.
  • Mac: the same totals apply to unified memory, so the 40–43 GB builds need a 64 GB machine. See running Llama 3 on a Mac.
  • No GPU: the file loads into system RAM, so 64 GB of RAM covers Q4_K_M.
  • Disk: the download size in the table, once per build you pull.

This site has not measured generation speed for this model, so none is quoted. For what else fits each tier see best LLM for 24 GB of VRAM, best LLM for 32 GB of VRAM and the Ollama RAM and VRAM table; splitting a model across cards is covered in the Ollama multi-GPU guide, and the build names in quantization explained.

Llama 3 vs 3.1 vs 3.3 70B sizes on Ollama

All three 70B generations have the same 70,553,706,496 parameters, so they take the same memory at the same quantisation. The reason llama3:70b is 40 GB and the other two are 43 GB is that the Llama 3 default tag is the older q4_0 format while the 3.1 and 3.3 defaults are Q4_K_M.

Ollama tagDownloadQuantisationContextModel released
llama3:70b40 GBq4_08K18 April 2024
llama3:70b-instruct-q4_K_M43 GBQ4_K_M8K18 April 2024
llama3.1:70b43 GBQ4_K_M128K23 July 2024
llama3.3:70b43 GBQ4_K_M128KLate 2024

Sizes are from the Ollama tag lists for llama3, llama3.1 and llama3.3. Note that llama3 with no tag is the 8B model (4.7 GB), not the 70B.

Model specifications

SpecLlama 3 70B
DeveloperMeta
Release date18 April 2024
Parameters70,553,706,496
Context length8K (8,192 tokens)
ArchitectureDense transformer with grouped-query attention: 80 layers, width 8192, 64 attention heads, 8 key-value heads
Vocabulary128K tokens
Training dataOver 15 trillion tokens of publicly available data
Knowledge cutoffDecember 2023
VariantsBase (pretrained) and Instruct; input and output are text only
Intended languageEnglish
LicenceMeta Llama 3 Community License

From Meta's Llama 3 model card; the layer and head counts are from the model's configuration file.

Benchmarks published by Meta

Llama 3 70B Instruct scores 82.0 on MMLU (5-shot) on Meta's model card; the base model scores 79.5. These are Meta's own results from its internal evaluation library. The card compares Llama 3 only with Llama 2; it contains no figures for any other company's models.

Benchmark (instruct models)Llama 3 70BLlama 2 70B
MMLU (5-shot)82.052.9
GPQA (0-shot)39.521.0
HumanEval (0-shot)81.725.6
GSM-8K (8-shot, CoT)93.057.5
MATH (4-shot, CoT)50.411.6

For the base model the card reports MMLU 79.5, ARC-Challenge 93.0, Winogrande 83.1, CommonSenseQA 83.8, BIG-Bench Hard 81.3, TriviaQA-Wiki 89.7 and DROP 79.7. Earlier versions of this page quoted an MMLU of 79.2 and a HumanEval of 67.0; neither figure is on Meta's card and both have been removed. What the benchmark names mean is covered in the guide to AI benchmarks.

Installation

With Ollama installed, pull the tag that matches your memory. Every tag below is on ollama.com/library/llama3/tags.

# Default 70B build: q4_0, 40 GB — for 48 GB of VRAM
ollama pull llama3:70b
ollama run llama3:70b

# Q4_K_M, 43 GB
ollama pull llama3:70b-instruct-q4_K_M

# Q2_K, 26 GB — for one 32 GB card
ollama pull llama3:70b-instruct-q2_K

For a quantisation Ollama does not carry, such as the 19 GB IQ2_XXS file, download the GGUF from the bartowski repository linked above and load it in llama.cpp. The original weights are on Hugging Face; the repository is gated, so you accept Meta's licence before downloading.

Unless you need this exact model, pull llama3.3:70b instead. It is 43 GB, runs on the same hardware and has a 128K context.

Licence

Llama 3 is released under the Meta Llama 3 Community License. Meta describes it as a custom commercial licence. It is not an open-source licence such as MIT or Apache 2.0. The terms that matter in practice:

  • Commercial use is permitted, along with modification and redistribution.
  • If you distribute the model or a product built on it, you must include the licence and display “Built with Meta Llama 3”.
  • A model you train or fine-tune from it and distribute must have a name beginning with “Llama 3”.
  • A company with more than 700 million monthly active users must request a separate licence from Meta.
  • Use is also subject to Meta's Acceptable Use Policy.

This is a summary, not legal advice; read the licence before relying on it.

What has replaced it

Llama 3 70B has been superseded twice at the same size, and the 8K context is the main reason to move on:

  • Llama 3.1 70B — released 23 July 2024 with a 128K context. Meta's card for it lists the longer context window, multilingual input and output, and tool integration as the new capabilities.
  • Llama 3.3 70B — the most recent 70B text model in the Llama 3 line, also 128K context. It is the one to pull today.
  • Llama 4 Scout — Meta's next generation accepts images as well as text. The Ollama llama4:scout tag is 67 GB, so it needs more memory than any 70B build at 4-bit.

If you have 48 GB of VRAM and want the strongest model that fits, rather than this model specifically, the best Ollama models list is kept current. Llama 3 70B is still worth keeping if you already have it and your prompts fit in 8K tokens.

Frequently asked questions

How big is Llama 3 70B?

The default Ollama tag, llama3:70b, is 40 GB. Other builds range from 26 GB at Q2_K to 75 GB at Q8_0, and the full-precision fp16 build is 141 GB.

How much VRAM does Llama 3 70B use at Q4_K_M?

About 46 GB with the full 8K context: a 42.5 GB file, 2.7 GB of context cache and roughly 1 GB of buffers. That is calculated from the file size and the model configuration, not measured.

What is the official MMLU score?

82.0 for the instruct model and 79.5 for the base model, both 5-shot, from Meta's model card.

When was Llama 3 70B released?

18 April 2024, together with Llama 3 8B.

Should I use Llama 3 70B or Llama 3.3 70B?

Llama 3.3 70B. It has the same parameter count, so it runs on the same hardware, and its context window is 128K instead of 8K.

What should I use to serve it to many users at once?

Ollama is built for one machine and a handful of requests. For many simultaneous requests the usual choice is a batching server such as vLLM running the original weights, which at 141 GB in 16-bit precision means two 80 GB cards. The cost is the hardware; the model itself is free to download. This site has not run that configuration and has no throughput figures for it.

Sources

Reading now
Join the discussion

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📅 Published: September 25, 2025🔄 Last Updated: September 29, 2026✓ Manually Reviewed

Related Guides

Continue your local AI journey with these comprehensive guides

More on AI Models Directory
See the full AI Models Directory guide.
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Found your model? Now build something with it.

25 hands-on courses — RAG, agents, fine-tuning — all running locally. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators