Llama 3 70B:
Technical Analysis & Setup
Updated: September 29, 2026
Llama 3 70B is the larger of the two models Meta released on 18 April 2024. It has 70.6 billion parameters and an 8K context window, and the default Ollama build, llama3:70b, is a 40 GB download. That puts it on 48 GB of VRAM, or on one 32 GB card at 2-bit quantisation. It has been replaced by Llama 3.1 70B and Llama 3.3 70B, which are the same size with a 128K context. Everything on this page is copied from Meta's model card and licence, and from the file sizes published on Ollama and Hugging Face; this site has not benchmarked the model.
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Llama 3 70B hardware requirements
Llama 3 70B needs about 46 GB of memory at Q4_K_M with its full 8K context, so the standard setup is 48 GB of VRAM: two 24 GB cards or one 48 GB card. The 26 GB Q2_K build fits a single 32 GB card. A single 24 GB card holds only the 19 GB IQ2_XXS build, or runs a larger build with part of the model in system RAM. The arithmetic is parameters × bits per weight ÷ 8: 70.55B × 4.82 bits ÷ 8 = 42.5 GB at Q4_K_M.
File sizes are from bartowski/Meta-Llama-3-70B-Instruct-GGUF and Ollama's tag list. The context cache is arithmetic from the model's config.json: 80 layers × 8 key-value heads × 128 dimensions, stored twice at 2 bytes, is 328 KB per token, or 2.7 GB for 8,192 tokens. “Memory needed” is file + that cache + a 1 GB allowance for buffers.
| Quantisation | File | Ollama tag (listed size) | Memory needed (8K context) | Fits |
|---|---|---|---|---|
| IQ2_XXS | 19.10 GB | not on Ollama | ~22.8 GB | One 24 GB card (RTX 3090, RTX 4090), with heavy quality loss |
| Q2_K | 26.38 GB | 70b-instruct-q2_K (26GB) | ~30.1 GB | One 32 GB card (RTX 5090) |
| Q3_K_M | 34.27 GB | 70b-instruct-q3_K_M (34GB) | ~38.0 GB | 48 GB: two 24 GB cards or one 48 GB card |
| q4_0 | 40 GB (Ollama) | llama3:70b (40GB) | ~43.7 GB | 48 GB: two 24 GB cards or one 48 GB card |
| Q4_K_M | 42.52 GB | 70b-instruct-q4_K_M (43GB) | ~46.2 GB | 48 GB, with under 2 GB to spare |
| Q5_K_M | 49.95 GB | 70b-instruct-q5_K_M (50GB) | ~53.7 GB | 64 GB: two 32 GB cards or three 24 GB cards |
| Q6_K | 57.89 GB | 70b-instruct-q6_K (58GB) | ~61.6 GB | 64 GB, tight; 72 GB across three 24 GB cards |
| Q8_0 | 75 GB (Ollama) | 70b-instruct-q8_0 (75GB) | ~78.7 GB | 96 GB |
| fp16 | 141.12 GB | 70b-instruct-fp16 (141GB) | ~144.8 GB | Two 80 GB data-centre cards |
- One 24 GB card: the 40 GB default does not fit. Ollama will split the model between the card and system RAM, which works but is slower than running fully on the GPU; you need about 48 GB of combined memory.
- Mac: the same totals apply to unified memory, so the 40–43 GB builds need a 64 GB machine. See running Llama 3 on a Mac.
- No GPU: the file loads into system RAM, so 64 GB of RAM covers Q4_K_M.
- Disk: the download size in the table, once per build you pull.
This site has not measured generation speed for this model, so none is quoted. For what else fits each tier see best LLM for 24 GB of VRAM, best LLM for 32 GB of VRAM and the Ollama RAM and VRAM table; splitting a model across cards is covered in the Ollama multi-GPU guide, and the build names in quantization explained.
Llama 3 vs 3.1 vs 3.3 70B sizes on Ollama
All three 70B generations have the same 70,553,706,496 parameters, so they take the same memory at the same quantisation. The reason llama3:70b is 40 GB and the other two are 43 GB is that the Llama 3 default tag is the older q4_0 format while the 3.1 and 3.3 defaults are Q4_K_M.
| Ollama tag | Download | Quantisation | Context | Model released |
|---|---|---|---|---|
llama3:70b | 40 GB | q4_0 | 8K | 18 April 2024 |
llama3:70b-instruct-q4_K_M | 43 GB | Q4_K_M | 8K | 18 April 2024 |
llama3.1:70b | 43 GB | Q4_K_M | 128K | 23 July 2024 |
llama3.3:70b | 43 GB | Q4_K_M | 128K | Late 2024 |
Sizes are from the Ollama tag lists for llama3, llama3.1 and llama3.3. Note that llama3 with no tag is the 8B model (4.7 GB), not the 70B.
Model specifications
| Spec | Llama 3 70B |
|---|---|
| Developer | Meta |
| Release date | 18 April 2024 |
| Parameters | 70,553,706,496 |
| Context length | 8K (8,192 tokens) |
| Architecture | Dense transformer with grouped-query attention: 80 layers, width 8192, 64 attention heads, 8 key-value heads |
| Vocabulary | 128K tokens |
| Training data | Over 15 trillion tokens of publicly available data |
| Knowledge cutoff | December 2023 |
| Variants | Base (pretrained) and Instruct; input and output are text only |
| Intended language | English |
| Licence | Meta Llama 3 Community License |
From Meta's Llama 3 model card; the layer and head counts are from the model's configuration file.
Benchmarks published by Meta
Llama 3 70B Instruct scores 82.0 on MMLU (5-shot) on Meta's model card; the base model scores 79.5. These are Meta's own results from its internal evaluation library. The card compares Llama 3 only with Llama 2; it contains no figures for any other company's models.
| Benchmark (instruct models) | Llama 3 70B | Llama 2 70B |
|---|---|---|
| MMLU (5-shot) | 82.0 | 52.9 |
| GPQA (0-shot) | 39.5 | 21.0 |
| HumanEval (0-shot) | 81.7 | 25.6 |
| GSM-8K (8-shot, CoT) | 93.0 | 57.5 |
| MATH (4-shot, CoT) | 50.4 | 11.6 |
For the base model the card reports MMLU 79.5, ARC-Challenge 93.0, Winogrande 83.1, CommonSenseQA 83.8, BIG-Bench Hard 81.3, TriviaQA-Wiki 89.7 and DROP 79.7. Earlier versions of this page quoted an MMLU of 79.2 and a HumanEval of 67.0; neither figure is on Meta's card and both have been removed. What the benchmark names mean is covered in the guide to AI benchmarks.
Installation
With Ollama installed, pull the tag that matches your memory. Every tag below is on ollama.com/library/llama3/tags.
# Default 70B build: q4_0, 40 GB — for 48 GB of VRAM
ollama pull llama3:70b
ollama run llama3:70b
# Q4_K_M, 43 GB
ollama pull llama3:70b-instruct-q4_K_M
# Q2_K, 26 GB — for one 32 GB card
ollama pull llama3:70b-instruct-q2_KFor a quantisation Ollama does not carry, such as the 19 GB IQ2_XXS file, download the GGUF from the bartowski repository linked above and load it in llama.cpp. The original weights are on Hugging Face; the repository is gated, so you accept Meta's licence before downloading.
Unless you need this exact model, pull llama3.3:70b instead. It is 43 GB, runs on the same hardware and has a 128K context.
Licence
Llama 3 is released under the Meta Llama 3 Community License. Meta describes it as a custom commercial licence. It is not an open-source licence such as MIT or Apache 2.0. The terms that matter in practice:
- Commercial use is permitted, along with modification and redistribution.
- If you distribute the model or a product built on it, you must include the licence and display “Built with Meta Llama 3”.
- A model you train or fine-tune from it and distribute must have a name beginning with “Llama 3”.
- A company with more than 700 million monthly active users must request a separate licence from Meta.
- Use is also subject to Meta's Acceptable Use Policy.
This is a summary, not legal advice; read the licence before relying on it.
What has replaced it
Llama 3 70B has been superseded twice at the same size, and the 8K context is the main reason to move on:
- Llama 3.1 70B — released 23 July 2024 with a 128K context. Meta's card for it lists the longer context window, multilingual input and output, and tool integration as the new capabilities.
- Llama 3.3 70B — the most recent 70B text model in the Llama 3 line, also 128K context. It is the one to pull today.
- Llama 4 Scout — Meta's next generation accepts images as well as text. The Ollama
llama4:scouttag is 67 GB, so it needs more memory than any 70B build at 4-bit.
If you have 48 GB of VRAM and want the strongest model that fits, rather than this model specifically, the best Ollama models list is kept current. Llama 3 70B is still worth keeping if you already have it and your prompts fit in 8K tokens.
Frequently asked questions
How big is Llama 3 70B?
The default Ollama tag, llama3:70b, is 40 GB. Other builds range from 26 GB at Q2_K to 75 GB at Q8_0, and the full-precision fp16 build is 141 GB.
How much VRAM does Llama 3 70B use at Q4_K_M?
About 46 GB with the full 8K context: a 42.5 GB file, 2.7 GB of context cache and roughly 1 GB of buffers. That is calculated from the file size and the model configuration, not measured.
What is the official MMLU score?
82.0 for the instruct model and 79.5 for the base model, both 5-shot, from Meta's model card.
When was Llama 3 70B released?
18 April 2024, together with Llama 3 8B.
Should I use Llama 3 70B or Llama 3.3 70B?
Llama 3.3 70B. It has the same parameter count, so it runs on the same hardware, and its context window is 128K instead of 8K.
What should I use to serve it to many users at once?
Ollama is built for one machine and a handful of requests. For many simultaneous requests the usual choice is a batching server such as vLLM running the original weights, which at 141 GB in 16-bit precision means two 80 GB cards. The cost is the hardware; the model itself is free to download. This site has not run that configuration and has no throughput figures for it.
Sources
- Meta Llama 3 model card — release date, context, training data, benchmark tables
- Meta Llama 3 Community License — the licence text
- Meta's Llama 3 announcement
- Meta-Llama-3-70B-Instruct on Hugging Face — original weights
- Ollama tag list — every tag with its download size
- bartowski GGUF files — the file sizes in the hardware table
- Llama 3.1 model card and Llama 3.3 model card
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Written by the Local AI Master Team
The team behind Local AI Master
We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.
Related Guides
Continue your local AI journey with these comprehensive guides
- PILLARLocal AI Models Directory: Every Model Compared
- Alpaca 7B: Stanford\
- Amazon Chronos: Time Series Forecasting Models (Complete Guide)
- Anima 2.9B on 8GB: The Anime Model Taking Over Civitai
- Aquila 7B by BAAI: Chinese-English Bilingual (FlagAI)
- Baichuan2-13B: Chinese LLM | 59% CMMLU, Bilingual, Free License 2026
- Bark by Suno AI: Open-Source Text-to-Audio Generation Guide
- ChatGLM3-6B: Tsinghua Chinese AI | Code Interpreter, 6GB RAM 2026
- Claude 3 Opus Review: Benchmarks, Pricing & API Guide 2026
- Claude 3 Sonnet Review: Benchmarks, API Pricing & Alternatives 2026
Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide
No spam. Unsubscribe with one click.
Found your model? Now build something with it.
25 hands-on courses — RAG, agents, fine-tuning — all running locally. First chapter free, no card.