★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once

Llama 2 70B: architecture, file sizes and hardware requirements

Updated: September 29, 2026

Llama 2 70B has 80 layers, a hidden size of 8192, 64 attention heads with 8 key/value heads (grouped-query attention), a 32,000-token vocabulary and a 4,096-token context window. It holds 68.98 billion parameters, which is 138 GB at fp16 and 41 GB at Q4_K_M. No published quantisation fits a single 24 GB card; the usual home setup is two 24 GB cards or a 64 GB Mac.

Llama 2 70B is a July 2023 model and has been replaced. Llama 3.1 70B and Llama 3.3 70B need about the same memory and have a 128K context instead of 4K. Use this page if you are maintaining something built on Llama 2 or need its specifications.

How many layers does Llama 2 70B have?

80 layers, each 8192 wide. The values below are copied from the config.json shipped with the weights. Meta's own repositories are gated behind a licence form, so they were read from two ungated fp16 mirrors, TheBloke/Llama-2-70B-fp16 and NousResearch/Llama-2-70b-hf, which agree on every field.

FieldLlama 2 70BNote
Parameters68,976,653,312 (68.98B)Hugging Face safetensors count; marketed as 70B
Layers (num_hidden_layers)80Decoder-only transformer blocks
Hidden size8192Width of the residual stream
Attention heads64128 dimensions per head (8192 / 64)
Key/value heads8Grouped-query attention; the 7B and 13B do not use it
MLP intermediate size286723.5 x the hidden size
Vocabulary32,000SentencePiece BPE
Context length4,096 tokensmax_position_embeddings; Meta calls it 4k
Weights dtypefloat16138 GB on disk

Where the 68.98 billion parameters come from

The count follows from the config. In each layer the query and output projections are 8192 × 8192 and the key and value projections are 8192 × 1024 (8 key/value heads of 128 dimensions), which comes to 151.0 million attention weights. The MLP has three 8192 × 28672 matrices, 704.6 million weights. That is 855.6 million per layer and 68.45 billion across 80 layers. The input embedding and the output head are each 32,000 × 8192, adding 0.52 billion. The total, 68.98 billion, matches the 68,976,653,312 that Hugging Face reports for the safetensors files.

Grouped-query attention is the one architectural difference between the 70B and its smaller siblings. Meta's model card marks GQA as used only on the 70B, “for improved inference scalability”. Sharing 8 key/value heads across 64 query heads makes the KV cache an eighth of what full multi-head attention would need. The Llama 2 13B uses standard attention.

Llama 2 70B hardware requirements

Plan for about 44 GB of GPU or unified memory to run the Q4_K_M file with its full context: two 24 GB cards, one 48 GB card, or a 64 GB Mac. The smallest published file, Q2_K, is 29 GB and still does not fit a 24 GB card. The arithmetic is parameters × bits per weight: 68.98 billion × 16 bits ÷ 8 is 138 GB at fp16, and the 41.4 GB Q4_K_M file works out to 4.8 bits per weight.

File sizes are from Ollama's llama2 tag list and agree with the GGUF files in TheBloke/Llama-2-70B-Chat-GGUF. The memory column is the file plus 1.3 GB of KV cache for the full 4,096-token context (80 layers × 8 KV heads × 128 dimensions × 2 × 2 bytes is 328 KB per token) plus a 1 GB allowance for compute buffers. It is arithmetic, not a measurement.

QuantisationOllama tagFileMemory neededWhat fits it
Q2_Kllama2:70b-chat-q2_K29 GB~32 GBOne 32 GB RTX 5090 with nothing else on the card, and it is tight.
Q3_K_Mllama2:70b-chat-q3_K_M33 GB~36 GBTwo 24 GB cards, or one 48 GB card.
Q4_0 (Ollama default)llama2:70b39 GB~41 GBTwo 24 GB cards (RTX 3090 / 4090), one 48 GB card, or a 64 GB Mac.
Q4_K_Mllama2:70b-chat-q4_K_M41 GB~44 GBTwo 24 GB cards with about 4 GB to spare, one 48 GB card, or a 64 GB Mac.
Q5_K_Mllama2:70b-chat-q5_K_M49 GB~51 GBDoes not fit 48 GB. Two 32 GB cards or three 24 GB cards.
Q6_Kllama2:70b-chat-q6_K57 GB~59 GBTwo 32 GB cards, three 24 GB cards, or a 96 GB Mac.
Q8_0llama2:70b-chat-q8_073 GB~76 GBOne 80 GB or 96 GB workstation card, or four 24 GB cards.
fp16llama2:70b-chat-fp16138 GB~140 GBTwo 80 GB cards. Not a desktop configuration.

With one 24 GB card, llama.cpp and Ollama will put as many layers as fit on the GPU and run the rest from system RAM. That works with 64 GB of RAM, but the layers left on the CPU set the pace, so expect it to be slow. On a Mac, macOS limits how much unified memory the GPU may use, which is why 64 GB is the practical floor for the 4-bit files; the Mac memory pressure guide explains the limit.

If you have a single 24 GB or 32 GB card, a newer and smaller model is a better use of it than a 2-bit Llama 2 70B: see the best LLM for 24 GB of VRAM and the 32 GB tier. The VRAM calculator runs the same sums for other models, and quantization explained covers what the K-quant names mean.

Ollama tags and install

llama2:70b is the chat model at q4_0, a 39 GB download. It shares a digest with llama2:70b-chat and llama2:70b-chat-q4_0, so all three names pull the same file. The 70b-text tags are the pretrained completion model with no chat tuning. Every 70B tag is listed with a 4K context.

# Chat model, q4_0, 39 GB
ollama pull llama2:70b
ollama run llama2:70b

# Q4_K_M, 41 GB
ollama pull llama2:70b-chat-q4_K_M

# Smallest published file, Q2_K, 29 GB
ollama pull llama2:70b-chat-q2_K

# Pretrained completion model (not chat-tuned), 39 GB
ollama pull llama2:70b-text

The original weights are on Hugging Face at meta-llama/Llama-2-70b-chat-hf, behind Meta's licence form. Memory by model size for everything else on Ollama is in the Ollama system requirements guide.

Benchmarks Meta reported

These are Meta's numbers for the pretrained 70B, copied from the table in the Llama 2 model card and produced with Meta's internal evaluation library. This site has not run them. Most rows are averages over several benchmarks, as the labels say. On TruthfulQA the card reports 50.18 for the pretrained 70B and 64.14 for the chat model.

Benchmark groupLlama 2 70B
MMLU (5-shot)68.9
Code (average of HumanEval and MBPP pass@1)37.5
Commonsense reasoning (7-benchmark average)71.9
World knowledge (NaturalQuestions and TriviaQA, 5-shot)63.6
Reading comprehension (SQuAD, QuAC, BoolQ, 0-shot)69.4
Math (GSM8K 8-shot and MATH 4-shot average)35.2
BIG-Bench Hard51.2
AGI Eval54.2

Training details from the same card: 2.0 trillion pretraining tokens, trained between January and July 2023, pretraining data cut off in September 2022, and 1,720,320 GPU hours on A100-80GB hardware for the 70B. The full method is in the paper, Llama 2: Open Foundation and Fine-Tuned Chat Models (arXiv:2307.09288).

What replaced Llama 2 70B

Meta has shipped three generations since. If you have the hardware for Llama 2 70B you have the hardware for either 70B successor, and the 4K context is the main reason to move: it is small enough that a long document or a multi-file coding session will not fit. Sizes are from the Ollama tag lists linked under Sources.

ModelOllama tagDownloadContextNote
Llama 3.1 70Bllama3.1:70b43 GB128KMeta reports MMLU 79.3 for the pretrained model, against 68.9 for Llama 2 70B.
Llama 3.3 70Bllama3.3:70b43 GB128KReleased 6 December 2024. Same memory class as Llama 2 70B.
Llama 4 Scoutllama4:scout67 GB10M as listedMixture-of-experts; a larger download than a 70B at 4-bit.
Llama 2 70Bllama2:70b39 GB4KThis page.

Related pages from the same era: CodeLlama 70B, the code-tuned model built on Llama 2, and Mixtral 8x7B. For what to run today on two 24 GB cards or one, start with the best local models for a 24 GB GPU.

Frequently asked questions

What is the context length of Llama 2 70B?

4,096 tokens. The config sets max_position_embeddings to 4096, Meta's model card lists 4k for all three sizes, and Ollama lists every 70B tag at 4K. There is no long-context variant of Llama 2.

How big is Llama 2 70B on disk?

138 GB at fp16, 73 GB at Q8_0, 41 GB at Q4_K_M, 39 GB at the q4_0 that Ollama pulls by default, and 29 GB at Q2_K. People often quote 140 GB for the full-precision weights; that is 70 billion × 2 bytes, and the exact figure is lower because the model has 68.98 billion parameters.

Can Llama 2 70B run on one RTX 4090?

Not entirely on the card. The smallest file is 29 GB and the card has 24 GB. It will run with part of the model in system RAM, slowly. Two 24 GB cards hold the 4-bit files with the full context.

What licence is Llama 2 under?

The Llama 2 Community License, which Meta's model card describes as a custom commercial licence, together with an Acceptable Use Policy. It is not an open-source licence in the OSI sense. Read it before you ship anything built on the model.

Is there any reason to start a new project on Llama 2 70B?

In my view, no. Llama 3.1 70B and Llama 3.3 70B need about the same memory and have 32 times the context, and Meta reports a higher MMLU score for Llama 3.1 70B (79.3 against 68.9). Llama 2 70B is worth keeping if you depend on a fine-tune built on it or need to reproduce older results.

Sources

Learn AI in the Right Order

Structured courses with hands-on projects and local-first workflows that reduce API dependency where they fit.

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📅 Published: 2023-07-18🔄 Last Updated: September 29, 2026✓ Manually Reviewed
Reading now
Join the discussion
More on AI Models Directory
See the full AI Models Directory guide.
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Found your model? Now build something with it.

25 hands-on courses — RAG, agents, fine-tuning — all running locally. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators