Copy Ollama Models to an Offline PC, No Re-Download
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Ollama’s running. Here’s what to build with it. Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.
Short answer: copy two folders — blobs/ and manifests/ — out of the Ollama models directory, put them in the same place on the offline machine, fix ownership, restart the service. No re-download, no internet, no registry. Ollama's manifests reference weights by content hash rather than by URL or absolute path, so a straight file copy is a complete transfer.
That is the whole trick, and it is why ollama pull hanging on a disconnected workstation is an annoyance rather than a wall. The rest of this page is the exact procedure, the three ways people get it wrong, how to prove a 40GB blob survived the USB stick, and what to do on a machine where you cannot even run the installer.
Everything on this page is either documented by Ollama — the storage paths, the OLLAMA_MODELS variable, the chown requirement on Linux, the import syntax — or falls straight out of how the blob store is built, which is content-addressed in the same way an OCI registry is. Sources are listed at the bottom and attributed inline. Where a claim would need a benchmark or a packet capture to stand up, this page says so rather than asserting it.
The Short Version
Five commands, if you are in a hurry and both machines are Linux.
# On the connected machine
sudo tar --exclude='*-partial*' -czf ollama-models.tar.gz \
-C /usr/share/ollama/.ollama models
# Sneakernet the tarball across
# On the air-gapped machine
sudo systemctl stop ollama
sudo tar -xzf ollama-models.tar.gz -C /usr/share/ollama/.ollama
sudo chown -R ollama:ollama /usr/share/ollama/.ollama/models
sudo systemctl start ollama && ollama list
If ollama list shows your models, you are done. If it does not, the what goes wrong section covers the three causes in order of how often each one is the real one.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
Where the Files Live
Ollama's FAQ publishes the model directory per platform. There is no hidden second location.
| OS | Models directory (per Ollama FAQ) |
|---|---|
| macOS | ~/.ollama/models |
| Linux | /usr/share/ollama/.ollama/models |
| Windows | C:\Users\%username%\.ollama\models |
Two footnotes worth knowing. The Windows docs describe the same location as %HOMEPATH%\.ollama — same path, different notation, and searchers hit both spellings. And the Linux path is the service account's home, not yours: if you installed with the official script, Ollama runs as the ollama user, which is the source of most permission failures later on.
To relocate the directory entirely — onto an encrypted volume, a bigger disk, or a share — set OLLAMA_MODELS. The FAQ is explicit that on Linux with the standard installer you must then give the service account access:
sudo chown -R ollama:ollama /mnt/models
On Windows the docs say to set OLLAMA_MODELS through the Settings / Control Panel environment-variable editor and restart Ollama for the change to take effect.
The single most common mistake with OLLAMA_MODELS is setting it in the wrong process. The ollama CLI is a thin client that talks to the running daemon over HTTP; it does not read models off disk itself. Exporting the variable in the shell where you type ollama list therefore changes nothing — you can point it at a completely empty directory and ollama list will still return the daemon's normal model list, because the daemon never saw your export. The variable has to be set for the server process (the systemd unit, the Windows service, or the shell you run ollama serve in).
What the Layout Actually Is
Inside the models directory there are exactly two subdirectories: blobs/ and manifests/. You need both. Copying one without the other is the classic failure.
models/
├── blobs/
│ ├── sha256-735af2139dc652bf01112746474883d79a52fa1c19038265d363e3d42556f7a2
│ ├── sha256-4b19ac7dd2fb1ab2f2818b73454c5a9128ca39875a8fcf686a6b1c36100a0d68
│ └── ...
└── manifests/
└── registry.ollama.ai/
└── library/
├── gemma3/270m
├── qwen3/4b
└── llama3.2/latest
blobs/ holds every actual file — weights, prompt template, licence text, parameter defaults — each named sha256-<64 hex chars>, where the hex is the SHA-256 of the file's own contents. That is content addressing, and it is what makes the transfer safe: the same blob is shared by every model that uses it, and no blob knows or cares which machine it is on.
manifests/registry.ollama.ai/library/<model>/<tag> is a small JSON file — one per model tag, with no file extension. It is an OCI-style manifest listing the layers that make up that model. This is the shape of one, with the digests truncated for readability:
{
"schemaVersion": 2,
"mediaType": "application/vnd.docker.distribution.manifest.v2+json",
"config": { "digest": "sha256:74156d92caf6…", "size": 490 },
"layers": [
{ "mediaType": "application/vnd.ollama.image.model", "digest": "sha256:735af2139dc6…", "size": 291545472 },
{ "mediaType": "application/vnd.ollama.image.template", "digest": "sha256:4b19ac7dd2fb…", "size": 476 },
{ "mediaType": "application/vnd.ollama.image.license", "digest": "sha256:3e2c24001f9e…", "size": 8431 },
{ "mediaType": "application/vnd.ollama.image.params", "digest": "sha256:339e884a40f6…", "size": 61 }
]
}
Read it as a shopping list. Strip the sha256: prefix from each digest, prepend sha256-, and you have the exact blob filenames that model needs — typically four or five files: the weights, the prompt template, the licence text and the parameter defaults. That is how you move one 40GB model instead of the entire 300GB cache.
Two things to leave behind. Files with -partial in the name are interrupted downloads (sha256-0aa0872a…-partial, -partial-0, -partial-1, and so on). They are useless on the target and can be as large as whatever fraction of the model had downloaded before the pull died — often several gigabytes. Exclude them. Second, if you are copying selectively, do not copy a manifest whose blobs you skipped; you will get a named model that fails to load.
Transfer: Linux to Linux
The whole-directory copy is the one to use unless size forces your hand.
1. Find what you have and how big it is.
ollama list
sudo du -sh /usr/share/ollama/.ollama/models
2. Package it, excluding partial downloads.
sudo tar --exclude='*-partial*' -czf ~/ollama-models.tar.gz \
-C /usr/share/ollama/.ollama models
Model weights are already compressed; -z buys you very little and costs real time on 40GB. Drop it (-cf, plain tar) if the media has room — the transfer will finish considerably sooner.
3. Move the media. Nothing clever here. This is the step your security policy governs, not us.
4. Stop the service before writing into its directory.
sudo systemctl stop ollama
5. Unpack into place and hand it to the service account.
sudo tar -xf ~/ollama-models.tar.gz -C /usr/share/ollama/.ollama
sudo chown -R ollama:ollama /usr/share/ollama/.ollama/models
sudo systemctl start ollama
The chown is not optional and it is not paranoia — the Ollama FAQ calls it out for exactly this reason. If you extracted as root, every file is root-owned and the ollama service account cannot read a byte of it.
6. Confirm.
ollama list
ollama run gemma3:270m "Reply with the single word OK."
That second command is the one that matters. ollama list only proves the manifests parsed; a generation proves the blobs are intact and the runner can load them. Use the smallest model you transferred for this check — a sub-gigabyte model loads in seconds, and a successful generation on a models directory that never saw a network pull is the proof the whole procedure worked.
Selective transfer, if you only want one model, is the same procedure with a smaller tar:
cd /usr/share/ollama/.ollama/models
# read the manifest to get the blob list
cat manifests/registry.ollama.ai/library/qwen3/4b
# then tar that manifest plus each sha256-<hex> blob it names
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
Transfer: Windows to Windows
Same two folders. robocopy handles the long paths and the resume-on-failure that copy does not.
# On the connected machine — mirror to removable media, skipping partials
robocopy "$env:USERPROFILE\.ollama\models" "E:\ollama-models" /E /XF *-partial*
Quit Ollama from the taskbar on the target machine first, then:
robocopy "E:\ollama-models" "$env:USERPROFILE\.ollama\models" /E
Start Ollama again and run ollama list. Windows has no service-account ownership problem equivalent to the Linux ollama user — if the files are under your own profile you can read them — so the failure mode here is almost always an incomplete copy or a wrong path. robocopy prints a summary table at the end; a non-zero Failed column is your answer.
If you are copying to a different drive because C: is small, set OLLAMA_MODELS to the new path in the environment-variable editor and restart Ollama, as the Windows docs describe. If Ollama is not installed on the target at all, skip to machines with no installer. Our Ollama on Windows install guide covers the normal connected path.
Verify the Transfer
Every blob filename is its own checksum, so verification needs no manifest, no tool, and no trust in the media.
That is what content addressing means: hashing a blob named sha256-005f95c74751…acaf5 returns 005f95c74751…acaf5. Filename equals content hash, always — the same digests the manifest lists under layers. Nothing about that is machine-specific, which is precisely why the transfer works.
On Linux, verify the whole directory in one pass:
cd /usr/share/ollama/.ollama/models/blobs
for f in sha256-*; do
case "$f" in *partial*) continue;; esac
actual=$(sha256sum "$f" | cut -d' ' -f1)
[ "$actual" = "${f#sha256-}" ] || echo "CORRUPT: $f"
done
Silence means every blob is byte-identical to the source. On Windows:
Get-ChildItem "$env:USERPROFILE\.ollama\models\blobs" -Filter sha256-* |
Where-Object { $_.Name -notmatch 'partial' } |
ForEach-Object {
$h = (Get-FileHash $_.FullName -Algorithm SHA256).Hash.ToLower()
if ($h -ne $_.Name.Substring(7)) { "CORRUPT: $($_.Name)" }
}
Run this. Sneakernet media is the threat model in an air-gapped environment, and a silently truncated 20GB blob produces a model that loads and then generates garbage — which is a far worse outcome than a clean failure. For the wider control set around this (mirrors, signing, audit trail), our air-gapped AI deployment guide is the policy-layer companion to this file-level procedure.
Machines With No Installer
On Windows there is an official CLI-only archive, so administrator rights and MSI approval are not blockers.
The Ollama Windows documentation publishes a standalone zip "containing only the Ollama CLI and GPU library dependencies":
| Archive | For |
|---|---|
ollama-windows-amd64.zip | Base CLI plus NVIDIA GPU dependencies |
ollama-windows-amd64-rocm.zip | AMD ROCm acceleration |
ollama-windows-amd64-mlx.zip | MLX/CUDA variant |
Unzip to a directory you can write to, then:
$env:OLLAMA_MODELS = "D:\ollama\models"
.\ollama.exe serve
In a second terminal, .\ollama.exe list. Remember the client/server rule from earlier: OLLAMA_MODELS matters in the window running serve.
Two notes straight from the docs. If you are upgrading from a prior version, remove the old directories first. And the stated requirements are Windows 10 22H2 or newer, with NVIDIA driver 551.61 or newer for NVIDIA cards, or an AMD ROCm v7 / HIP7-capable driver stack (or a Vulkan-capable Radeon driver) on AMD. Driver installation may itself need admin rights on a locked-down box — that is the part this approach does not solve, and you should check it before promising anyone a delivery date. If no-admin is your actual constraint rather than no-network, that is a different problem with different answers.
Fallback: Import a Raw GGUF
If you would rather move a single .gguf file than a blob tree, Ollama will adopt it with a two-line Modelfile.
Per Ollama's import documentation, create a file named Modelfile:
FROM /path/to/model.gguf
Then:
ollama create my-model
The same doc covers Safetensors directories (FROM /path/to/safetensors/directory, or FROM . when the Modelfile sits alongside them) and adapters via ADAPTER /path/to/file.gguf — with the warning that an adapter must be created against the same base model you name in FROM, "otherwise you will get erratic results."
This path is genuinely useful when the GGUF came from somewhere other than Ollama's registry, or when the model was never in Ollama to begin with. The cost is that you get none of the registry model's prompt template, stop tokens, or parameter defaults — those live in separate blob layers on the blob path (look back at the manifest: image.template and image.params are their own layers), and a bare FROM x.gguf import ships without them. You will usually have to write the TEMPLATE and PARAMETER lines yourself; our Modelfile guide has the syntax. For a straight machine-to-machine move of something you already pulled, the blob copy is strictly better.
The Outbound-Traffic Question
Ollama's own FAQ confirms the desktop builds check for updates: "Ollama on macOS and Windows will automatically download updates." You apply it by clicking the taskbar or menubar item and then "Restart to update". On Linux, upgrading means re-running the install script — there is no background updater.
On a disconnected machine that check simply fails, but "it fails" is not the same as "it does not happen," and an audit that watches egress will log the attempt. Two practical responses:
- Run the CLI-only build on locked-down hosts. The standalone zip is the CLI and GPU libraries; there is no menubar updater in it.
- Enforce it at the network, not in the app. On an air-gapped site the firewall already denies the egress. That control is auditable; a client-side toggle is not.
We are not going to publish a list of endpoints to block, because we have not packet-captured any Ollama build and stale claims about what a binary connects to are worse than no claim at all. Point your own monitoring at it, and if your environment already runs an internal mirror, the air-gapped deployment guide covers that pattern properly. Once models are on the box, inference itself is fully local — what running AI offline actually means covers that side.
What Goes Wrong
Ordered by how often each one is the real cause.
1. ollama list is empty or missing the model you copied.
Almost always the client/server confusion: OLLAMA_MODELS set in your shell instead of for the daemon. Check what the server is actually using — on Linux, sudo systemctl show ollama -p Environment; on Windows, the service environment, not your terminal. Second most likely: you copied blobs/ but not manifests/, so there is no name to list. Third: you never restarted the service after copying.
2. Model listed but fails to load, or produces garbage. Run the checksum loop from verify. A truncated blob from a full USB stick or an interrupted copy is the classic cause, and it does not announce itself.
3. permission denied in the server log on Linux.
You extracted as root into the ollama user's directory. sudo chown -R ollama:ollama /usr/share/ollama/.ollama/models and restart. Check with ls -l — if you see root root on the blobs, that is the whole story.
4. Disk fills up mid-copy.
Usually -partial files that you meant to exclude — a pull interrupted at 80% of a 40GB model leaves 30GB behind. Check the source before packaging: ls -la models/blobs | grep partial.
5. Blobs seem to vanish after a restart.
Ollama exposes an OLLAMA_NOPRUNE setting, described in ollama serve --help as "Do not prune model blobs on startup". Whether a given release actually prunes orphaned blobs — ones whose manifest is missing — is the subject of a lot of forum folklore and we have not verified it on any version, so we are not going to tell you it does or does not. What we can say is that the setting costs nothing: if you are copying blobs before their manifests, or scripting a staged transfer, export OLLAMA_NOPRUNE=1 for the first startup and the question does not arise.
6. ollama pull hangs forever instead of failing.
That is the symptom that brought you here, and it is expected on a box with no route out — the client waits on a connection that will never be refused because nothing is there to refuse it. It is not a bug and there is no timeout worth tuning. Use the file transfer.
Limits of This Method
Four things this procedure does not do, stated plainly:
- It does not move Ollama itself. The binary and service still have to get onto the target — installer, standalone zip, or your organisation's package mirror.
- It does not carry your chat history, or anything outside
models/. Only weights and model metadata move. - It is not cross-architecture magic for the runtime. The blobs are portable; the Ollama build and GPU drivers on the target must match that machine's hardware. A model that ran on a 24GB card will still not fit an 8GB one — check the requirement first with our 8GB VRAM model picks or the RAM and VRAM table.
- It assumes the current directory layout. The
blobs/manifestssplit has been stable for a long time and Ollama's own FAQ documents the paths, but nothing guarantees it forever. Before trusting any blog post about this — including this one — open a manifest on your own machine and check it still looks like the example above.
If your air-gapped host will also be serving other people, lock the endpoint down before you hand out the address: our securing Ollama guide covers binding, auth and exposure.
Sources
- Ollama FAQ — model storage paths per OS,
OLLAMA_MODELS, thechown ollama:ollamainstruction, and automatic updates on macOS/Windows (fetched August 2026) - Ollama Windows documentation — standalone
ollama-windows-amd64.zip/-rocm/-mlxarchives, driver requirements,OLLAMA_MODELSon Windows - Ollama import documentation —
FROMsyntax for GGUF, Safetensors and adapters ollama serve --help— the environment-variable list, includingOLLAMA_NOPRUNE- The blob-filename/SHA-256 equivalence and the manifest layer list are properties of the content-addressed store itself, readable from any manifest file on your own machine
FAQ
Ollama’s running. Here’s what to build with it.
Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.
Stop piecing Ollama together from blog posts
Ollama Mastery is 15 chapters end to end — install, model choice, Modelfiles, GPU offload, the API, and the 20 errors that actually happen. Plus 24 more courses.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARBest Ollama Models 2026: 15 Ranked (Coding, Reasoning, Chat)
- AI on Steam Deck: Run Local LLMs with Ollama on SteamOS
- Air-Gapped AI Deployment: Install Ollama With No Internet
- Best Free Local AI Models to Run With Ollama (No API Key)
- Best Ollama Embedding Models Compared for Local RAG
- Best Ollama Models for 8GB RAM 2026: 12 Tested Local Picks
- Best Ollama Models for AI Agents 2026: Ranked by Tool Use
- Best Uncensored Local LLMs: Abliterated Ollama Models
- Browser-Use + Ollama: A Local Web-Browsing Agent
- Build a Local AI Slack & Discord Bot with Ollama + Python
Comments (0)
No comments yet. Be the first to share your thoughts!