Llama 2 70B: architecture, file sizes and hardware requirements
Updated: September 29, 2026
Llama 2 70B has 80 layers, a hidden size of 8192, 64 attention heads with 8 key/value heads (grouped-query attention), a 32,000-token vocabulary and a 4,096-token context window. It holds 68.98 billion parameters, which is 138 GB at fp16 and 41 GB at Q4_K_M. No published quantisation fits a single 24 GB card; the usual home setup is two 24 GB cards or a 64 GB Mac.
Llama 2 70B is a July 2023 model and has been replaced. Llama 3.1 70B and Llama 3.3 70B need about the same memory and have a 128K context instead of 4K. Use this page if you are maintaining something built on Llama 2 or need its specifications.
How many layers does Llama 2 70B have?
80 layers, each 8192 wide. The values below are copied from the config.json shipped with the weights. Meta's own repositories are gated behind a licence form, so they were read from two ungated fp16 mirrors, TheBloke/Llama-2-70B-fp16 and NousResearch/Llama-2-70b-hf, which agree on every field.
| Field | Llama 2 70B | Note |
|---|---|---|
| Parameters | 68,976,653,312 (68.98B) | Hugging Face safetensors count; marketed as 70B |
| Layers (num_hidden_layers) | 80 | Decoder-only transformer blocks |
| Hidden size | 8192 | Width of the residual stream |
| Attention heads | 64 | 128 dimensions per head (8192 / 64) |
| Key/value heads | 8 | Grouped-query attention; the 7B and 13B do not use it |
| MLP intermediate size | 28672 | 3.5 x the hidden size |
| Vocabulary | 32,000 | SentencePiece BPE |
| Context length | 4,096 tokens | max_position_embeddings; Meta calls it 4k |
| Weights dtype | float16 | 138 GB on disk |
Where the 68.98 billion parameters come from
The count follows from the config. In each layer the query and output projections are 8192 × 8192 and the key and value projections are 8192 × 1024 (8 key/value heads of 128 dimensions), which comes to 151.0 million attention weights. The MLP has three 8192 × 28672 matrices, 704.6 million weights. That is 855.6 million per layer and 68.45 billion across 80 layers. The input embedding and the output head are each 32,000 × 8192, adding 0.52 billion. The total, 68.98 billion, matches the 68,976,653,312 that Hugging Face reports for the safetensors files.
Grouped-query attention is the one architectural difference between the 70B and its smaller siblings. Meta's model card marks GQA as used only on the 70B, “for improved inference scalability”. Sharing 8 key/value heads across 64 query heads makes the KV cache an eighth of what full multi-head attention would need. The Llama 2 13B uses standard attention.
Llama 2 70B hardware requirements
Plan for about 44 GB of GPU or unified memory to run the Q4_K_M file with its full context: two 24 GB cards, one 48 GB card, or a 64 GB Mac. The smallest published file, Q2_K, is 29 GB and still does not fit a 24 GB card. The arithmetic is parameters × bits per weight: 68.98 billion × 16 bits ÷ 8 is 138 GB at fp16, and the 41.4 GB Q4_K_M file works out to 4.8 bits per weight.
File sizes are from Ollama's llama2 tag list and agree with the GGUF files in TheBloke/Llama-2-70B-Chat-GGUF. The memory column is the file plus 1.3 GB of KV cache for the full 4,096-token context (80 layers × 8 KV heads × 128 dimensions × 2 × 2 bytes is 328 KB per token) plus a 1 GB allowance for compute buffers. It is arithmetic, not a measurement.
| Quantisation | Ollama tag | File | Memory needed | What fits it |
|---|---|---|---|---|
| Q2_K | llama2:70b-chat-q2_K | 29 GB | ~32 GB | One 32 GB RTX 5090 with nothing else on the card, and it is tight. |
| Q3_K_M | llama2:70b-chat-q3_K_M | 33 GB | ~36 GB | Two 24 GB cards, or one 48 GB card. |
| Q4_0 (Ollama default) | llama2:70b | 39 GB | ~41 GB | Two 24 GB cards (RTX 3090 / 4090), one 48 GB card, or a 64 GB Mac. |
| Q4_K_M | llama2:70b-chat-q4_K_M | 41 GB | ~44 GB | Two 24 GB cards with about 4 GB to spare, one 48 GB card, or a 64 GB Mac. |
| Q5_K_M | llama2:70b-chat-q5_K_M | 49 GB | ~51 GB | Does not fit 48 GB. Two 32 GB cards or three 24 GB cards. |
| Q6_K | llama2:70b-chat-q6_K | 57 GB | ~59 GB | Two 32 GB cards, three 24 GB cards, or a 96 GB Mac. |
| Q8_0 | llama2:70b-chat-q8_0 | 73 GB | ~76 GB | One 80 GB or 96 GB workstation card, or four 24 GB cards. |
| fp16 | llama2:70b-chat-fp16 | 138 GB | ~140 GB | Two 80 GB cards. Not a desktop configuration. |
With one 24 GB card, llama.cpp and Ollama will put as many layers as fit on the GPU and run the rest from system RAM. That works with 64 GB of RAM, but the layers left on the CPU set the pace, so expect it to be slow. On a Mac, macOS limits how much unified memory the GPU may use, which is why 64 GB is the practical floor for the 4-bit files; the Mac memory pressure guide explains the limit.
If you have a single 24 GB or 32 GB card, a newer and smaller model is a better use of it than a 2-bit Llama 2 70B: see the best LLM for 24 GB of VRAM and the 32 GB tier. The VRAM calculator runs the same sums for other models, and quantization explained covers what the K-quant names mean.
Ollama tags and install
llama2:70b is the chat model at q4_0, a 39 GB download. It shares a digest with llama2:70b-chat and llama2:70b-chat-q4_0, so all three names pull the same file. The 70b-text tags are the pretrained completion model with no chat tuning. Every 70B tag is listed with a 4K context.
# Chat model, q4_0, 39 GB
ollama pull llama2:70b
ollama run llama2:70b
# Q4_K_M, 41 GB
ollama pull llama2:70b-chat-q4_K_M
# Smallest published file, Q2_K, 29 GB
ollama pull llama2:70b-chat-q2_K
# Pretrained completion model (not chat-tuned), 39 GB
ollama pull llama2:70b-textThe original weights are on Hugging Face at meta-llama/Llama-2-70b-chat-hf, behind Meta's licence form. Memory by model size for everything else on Ollama is in the Ollama system requirements guide.
Benchmarks Meta reported
These are Meta's numbers for the pretrained 70B, copied from the table in the Llama 2 model card and produced with Meta's internal evaluation library. This site has not run them. Most rows are averages over several benchmarks, as the labels say. On TruthfulQA the card reports 50.18 for the pretrained 70B and 64.14 for the chat model.
| Benchmark group | Llama 2 70B |
|---|---|
| MMLU (5-shot) | 68.9 |
| Code (average of HumanEval and MBPP pass@1) | 37.5 |
| Commonsense reasoning (7-benchmark average) | 71.9 |
| World knowledge (NaturalQuestions and TriviaQA, 5-shot) | 63.6 |
| Reading comprehension (SQuAD, QuAC, BoolQ, 0-shot) | 69.4 |
| Math (GSM8K 8-shot and MATH 4-shot average) | 35.2 |
| BIG-Bench Hard | 51.2 |
| AGI Eval | 54.2 |
Training details from the same card: 2.0 trillion pretraining tokens, trained between January and July 2023, pretraining data cut off in September 2022, and 1,720,320 GPU hours on A100-80GB hardware for the 70B. The full method is in the paper, Llama 2: Open Foundation and Fine-Tuned Chat Models (arXiv:2307.09288).
What replaced Llama 2 70B
Meta has shipped three generations since. If you have the hardware for Llama 2 70B you have the hardware for either 70B successor, and the 4K context is the main reason to move: it is small enough that a long document or a multi-file coding session will not fit. Sizes are from the Ollama tag lists linked under Sources.
| Model | Ollama tag | Download | Context | Note |
|---|---|---|---|---|
| Llama 3.1 70B | llama3.1:70b | 43 GB | 128K | Meta reports MMLU 79.3 for the pretrained model, against 68.9 for Llama 2 70B. |
| Llama 3.3 70B | llama3.3:70b | 43 GB | 128K | Released 6 December 2024. Same memory class as Llama 2 70B. |
| Llama 4 Scout | llama4:scout | 67 GB | 10M as listed | Mixture-of-experts; a larger download than a 70B at 4-bit. |
| Llama 2 70B | llama2:70b | 39 GB | 4K | This page. |
Related pages from the same era: CodeLlama 70B, the code-tuned model built on Llama 2, and Mixtral 8x7B. For what to run today on two 24 GB cards or one, start with the best local models for a 24 GB GPU.
Frequently asked questions
What is the context length of Llama 2 70B?
4,096 tokens. The config sets max_position_embeddings to 4096, Meta's model card lists 4k for all three sizes, and Ollama lists every 70B tag at 4K. There is no long-context variant of Llama 2.
How big is Llama 2 70B on disk?
138 GB at fp16, 73 GB at Q8_0, 41 GB at Q4_K_M, 39 GB at the q4_0 that Ollama pulls by default, and 29 GB at Q2_K. People often quote 140 GB for the full-precision weights; that is 70 billion × 2 bytes, and the exact figure is lower because the model has 68.98 billion parameters.
Can Llama 2 70B run on one RTX 4090?
Not entirely on the card. The smallest file is 29 GB and the card has 24 GB. It will run with part of the model in system RAM, slowly. Two 24 GB cards hold the 4-bit files with the full context.
What licence is Llama 2 under?
The Llama 2 Community License, which Meta's model card describes as a custom commercial licence, together with an Acceptable Use Policy. It is not an open-source licence in the OSI sense. Read it before you ship anything built on the model.
Is there any reason to start a new project on Llama 2 70B?
In my view, no. Llama 3.1 70B and Llama 3.3 70B need about the same memory and have 32 times the context, and Meta reports a higher MMLU score for Llama 3.1 70B (79.3 against 68.9). Llama 2 70B is worth keeping if you depend on a fine-tune built on it or need to reproduce older results.
Sources
- Llama 2 model card — Meta: sizes, context, GQA, training data, benchmark table, licence
- Llama 2: Open Foundation and Fine-Tuned Chat Models — the paper, arXiv:2307.09288, 18 July 2023
- config.json (fp16 mirror) — layers, hidden size, heads, vocabulary, context
- Ollama llama2 tags — every 70B tag with its download size
- TheBloke/Llama-2-70B-Chat-GGUF — GGUF files for llama.cpp
- Ollama llama3.1 tags, llama3.3 tags and llama4 tags — successor sizes
Learn AI in the Right Order
Structured courses with hands-on projects and local-first workflows that reduce API dependency where they fit.
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Written by the Local AI Master Team
The team behind Local AI Master
We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.
Related Guides
Continue your local AI journey with these comprehensive guides
- PILLARLocal AI Models Directory: Every Model Compared
- Alpaca 7B: Stanford\
- Amazon Chronos: Time Series Forecasting Models (Complete Guide)
- Anima 2.9B on 8GB: The Anime Model Taking Over Civitai
- Aquila 7B by BAAI: Chinese-English Bilingual (FlagAI)
- Baichuan2-13B: Chinese LLM | 59% CMMLU, Bilingual, Free License 2026
- Bark by Suno AI: Open-Source Text-to-Audio Generation Guide
- ChatGLM3-6B: Tsinghua Chinese AI | Code Interpreter, 6GB RAM 2026
- Claude 3 Opus Review: Benchmarks, Pricing & API Guide 2026
- Claude 3 Sonnet Review: Benchmarks, API Pricing & Alternatives 2026
Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide
No spam. Unsubscribe with one click.
Found your model? Now build something with it.
25 hands-on courses — RAG, agents, fine-tuning — all running locally. First chapter free, no card.