Ollama Latest Version (v0.33.2) + Version History
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Ollama’s running. Here’s what to build with it. Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.
Published on April 10, 2026 · Updated August 30, 2026 — 18 min read
The latest Ollama version is v0.33.2, released August 27, 2026. Run ollama --version to see what you have — if it is anything below 0.32, you are missing the new interactive agent mode (the bare ollama command now launches an agent that can chat, write code, and search the web), and anything below 0.19 lacks the MLX Apple Silicon backend (roughly 2× faster on Macs) plus a long list of new models. This page tracks every major Ollama release so you can decide whether to update and what to expect when you do.
What you will find here:
- Current stable version and what shipped in it
- Full release timeline from v0.1.0 to present
- Breaking changes that might affect your setup
- How to check, update, and roll back versions
- Version comparison table with key features
Ollama moves fast. The project has shipped over 40 releases since its first public beta, and the pace has only accelerated. Some releases add model support. Others overhaul the inference engine or change API behavior. Knowing what changed — and what broke — saves you hours of debugging.
If you are setting up Ollama for the first time, start with our complete Ollama guide instead. This page is for people who already run Ollama and want to understand the release cadence.
Table of Contents
- How to Check Your Ollama Version
- Current Stable Release
- How to Update Ollama
- Complete Version Timeline
- Breaking Changes by Version
- Version Comparison Table
- How to Roll Back to a Previous Version
- Release Cadence and Roadmap
- Troubleshooting Update Issues
- FAQ
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
How to Check Your Ollama Version
Three ways to find your current version:
Command Line
# Primary method
ollama --version
# Output: ollama version is 0.32.1
# Alternative — the version subcommand
ollama version
API Endpoint
# If Ollama is running as a service
curl http://localhost:11434/api/version
# Output: {"version":"0.33.2"}
System-Specific Locations
# macOS — check the app bundle
mdls -name kMDItemVersion /Applications/Ollama.app
# Linux — check the binary directly
ollama --version
# Windows — PowerShell
ollama --version
# Or check: Settings → Apps → Ollama
If ollama --version returns nothing or errors out, your installation is older than v0.1.17 (the version that added the flag) or the binary is not on your PATH. See the troubleshooting section below.
Current Stable Release
Ollama v0.33.2 (August 27, 2026)
The current stable release is a small bug-fix patch on top of the 0.33 line:
- The desktop app follows the system appearance again, restoring dark mode
- The macOS app now hands off to an already-running instance instead of launching a second one
- The Claude Desktop proxy no longer interrupts in-flight requests when the model catalog updates
v0.33.1 (August 26) added MLX support for Qwen3.8 Flash Next, structured output in the MLX runner, and a fix for Metal GPU timeouts when loading models from slow storage.
What actually changed in v0.33.0 (August 21, 2026)
This is the release worth reading if you are skipping versions.
Claude Desktop as a first-class target. You can now configure Claude Desktop to use Ollama as a third-party gateway provider, so a local model backs the desktop client directly.
Prefill caching was substantially reworked — the part most likely to affect you if you run agents:
- A cancelled prefill now keeps every restore point it crossed, so a retry resumes where it stopped instead of starting over
- Resumed prefills no longer record restore points that fail to cover what they claim. On models with recurrent layers this previously forced a request that matched 46k of 47k tokens to reprocess from zero
- Claude Code's "tokens left" countdown system message is now disabled, because Ollama moved it to the front of the prompt and it invalidated the KV cache on every single request
The 0.32.6 – 0.32.15 patches, if you are coming from further back:
- v0.32.15 (Aug 19) — model metadata is cached between requests, roughly halving time-to-first-token (~995 ms → ~524 ms in Ollama's own benchmark); fixed a wedge after a mid-stream parser error
- v0.32.11 (Aug 14) —
ollama launch dshfor DeepSeek Harness,ollama launch musefor Meta's Muse Code; web search in the OpenAI-compatible Responses API - v0.32.10 (Aug 12) — ⚠️ behaviour change: models that do not set
repeat_penaltynow default to 1.0 (off) instead of 1.1. This matches other engines and speeds up speculative decoding, but an older model may start repeating itself — set the parameter per-model if it does - v0.32.9 (Aug 11) — NVIDIA Nemotron 3.5 Lightning (30B MoE, 3B active) and the Nemotron 3 architecture
- v0.32.8 (Aug 10) — Muse Glimmer on all platforms
- v0.32.6 (Aug 4) — ⚠️
/v1/chat/completionsstreaming now matches OpenAI's wire format exactly (roleon the first chunk only,finish_reasonon its own chunk). Truncated responses reportfinish_reason: "length"instead of"tool_calls". Experimental image generation was removed — stay on v0.32.5 if you need it.
Release notes above are summarised from Ollama's published GitHub releases, checked August 30, 2026.
The 0.31 line that preceded it was a speed release:
- v0.31.1 (Jun 30) made Gemma 4 nearly 90% faster on average on Apple Silicon via multi-token prediction
- v0.31.2 (Jul 6) enabled flash attention on older NVIDIA GPUs (compute capability 6.x) and improved iGPU offloading for vision models
Previous stable line: v0.30.x (May–June 2026)
The 0.30 line (which opened with v0.30.0 on May 13, 2026) pairs the new MLX engine on Apple Silicon with continued llama.cpp improvements, so it runs well across both Mac and NVIDIA/AMD hardware.
Highlights of the 0.30 series:
- Improved compatibility and performance via an updated llama.cpp engine, alongside the MLX backend on Apple Silicon
- Broader hardware support, including GGUF-based models from Hugging Face and your own fine-tuned GGUF models
- Faster performance on NVIDIA hardware
- Metal GPU offload for multimodal models on Apple Silicon (fixing earlier cases where vision models ran on CPU)
- Command A and the North family now run on Apple Silicon through the MLX engine (added in v0.30.10)
Recent patch notes:
- v0.30.10 (Jun 17) — Command A / North family on Apple Silicon MLX, llama.cpp engine bump, MLX build-artifact fixes
- v0.30.9 (Jun 15) — Cohere2Moe architecture support, LFM2 parser/render fixes
- v0.30.8 (Jun 12) —
ollama launchprovider-selection fix, improved prompt caching, more stable MLX inference
Note:
nomic-embed-textnow lowercases inputs to match the model card. If you previously relied on mixed-case embeddings, re-embed your corpus after updating.
# Update to the latest version
# macOS (Homebrew)
brew upgrade ollama
# Linux
curl -fsSL https://ollama.com/install.sh | sh
# Windows — download installer from ollama.com/download
How to Update Ollama
macOS
# Homebrew (recommended)
brew update && brew upgrade ollama
# Verify update
ollama --version
# If using the .app bundle, it auto-updates on launch.
# Force a manual check:
open -a Ollama
# Click the menu bar icon → Check for Updates
Linux
# Official install script (always fetches latest)
curl -fsSL https://ollama.com/install.sh | sh
# If you installed via snap
sudo snap refresh ollama
# systemd users — restart the service after update
sudo systemctl restart ollama
Windows
# Download latest installer
# https://ollama.com/download/windows
# Or via winget
winget upgrade Ollama.Ollama
# Restart the Ollama service
net stop ollama && net start ollama
Docker
# Pull latest image
docker pull ollama/ollama:latest
# Stop and remove old container
docker stop ollama && docker rm ollama
# Start with new image
docker run -d --gpus all -v ollama:/root/.ollama \
-p 11434:11434 --name ollama ollama/ollama:latest
Important: Your models survive updates. Ollama stores models in ~/.ollama/models (Linux/macOS) or C:\Users\<you>\.ollama\models (Windows). Updating the binary does not delete them. Docker users should mount a volume as shown above.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
Complete Version Timeline
2026 Releases
v0.32.x Series (July 2026)
- v0.32.1 (Jul 16): Improved Gemma 4 tool calling; fixed MLX cache leaks; load-timeout improvements.
- v0.32.6 (Aug 4): OpenAI-compatible streaming now matches OpenAI's wire format; experimental image generation removed.
- v0.32.9 (Aug 11): NVIDIA Nemotron 3.5 Lightning and the Nemotron 3 architecture.
- v0.32.10 (Aug 12):
repeat_penaltynow defaults to 1.0 (off) when a model does not set it. - v0.32.15 (Aug 19): Model metadata cached between requests — time-to-first-token roughly halved.
- v0.33.0 (Aug 21): Claude Desktop gateway support; prefill restore points reworked.
- v0.33.2 (Aug 27): System-appearance/dark-mode fix; macOS single-instance handoff. Latest stable.
- v0.32.0 (Jul 11): Major release — the bare
ollamacommand now launches an interactive agent that can chat, write code, search the web, and delegate tasks; the Codex App integration was renamed ChatGPT
v0.31.x Series (June–July 2026)
- v0.31.2 (Jul 6): Flash attention enabled on older NVIDIA GPUs (compute capability 6.x); improved iGPU offloading for vision models
- v0.31.1 (Jun 30): Gemma 4 nearly 90% faster on average on Apple Silicon via multi-token prediction
v0.30.x Series (May–June 2026)
- v0.30.10 (Jun 17): Command A and the North family run on Apple Silicon via the MLX engine; llama.cpp engine bump; MLX build fixes
- v0.30.9 (Jun 15): Cohere2Moe architecture support; LFM2 parser/render fixes
- v0.30.8 (Jun 12):
ollama launchprovider-selection fix; improved prompt caching; more stable MLX inference - v0.30.0 (May 13): MLX engine on Apple Silicon now augmented by an updated llama.cpp engine for broader hardware support; GGUF models from Hugging Face and your own fine-tuned GGUFs; faster on NVIDIA; Metal GPU offload for multimodal models. Behavior change:
nomic-embed-textnow lowercases inputs to match the model card.
v0.19–v0.23 Series (March–May 2026)
- v0.23.1 (May 5): Gemma 4 MTP speculative decoding on Macs
- v0.21.1 (Apr 22): cloud-model launch shortcuts (e.g.
ollama launch kimi --model kimi-k2.6:cloud) - v0.19 (March 27, 2026): MLX backend for Apple Silicon (preview). Ollama switched its Mac inference path from direct Metal calls to Apple's MLX framework to exploit unified memory — roughly 2× decode speed on supported chips (one published benchmark went from ~58 to ~112 tok/s). The MLX preview requires 32GB+ of unified memory; base 8GB/16GB Macs keep using the existing engine. (The MLX preview was announced on the Ollama blog on March 30.)
v0.12–v0.14 Series (Sept 2025–Jan 2026)
- v0.14.0 (Jan 10, 2026): continued engine and model-support updates
- v0.13.0 (Nov 19, 2025): new architecture and tooling support
- v0.12.0 (Sept 18, 2025): cloud models, web search, and coding-agent integrations expanded across this period
2025 Releases
Mid-2025 highlights
- Sept 23, 2025: enhanced model scheduling with more precise memory management
- Web search API (2025): a hosted web-search endpoint with a free tier for individuals, letting local models pull current information to reduce hallucinations
- May 15, 2025: new multimodal engine — first-class vision support for models like
llama4:scout,gemma3,qwen2.5vl, andmistral-small3.1
v0.4.x Series (late 2024–early 2025)
- v0.4.7: Apple M3 Ultra optimized Metal shaders
- v0.4.5:
ollama showcommand for model metadata inspection - v0.4.2: Multimodal model support (LLaVA, BakLLaVA)
- v0.4.0: Breaking change — Model storage format migrated from blob-based to content-addressable. First run after update triggers automatic migration (can take 5–15 minutes depending on model count).
v0.3.x Series (October 2024–February 2025)
- v0.3.12: GGUF v3 format support
- v0.3.9:
ollama createfrom Safetensors (no manual conversion needed) - v0.3.6: AMD ROCm 5.7 support, Radeon RX 7900 XTX validated
- v0.3.0: Breaking change — Modelfile syntax updated.
ADAPTERcommand replacedFROM ... ADAPTERpattern.
v0.2.x Series (May–September 2024)
- v0.2.8: GPU layer offloading with
num_gpuparameter - v0.2.5: OpenAI-compatible API endpoint at
/v1/chat/completions - v0.2.0: Custom model creation via Modelfiles. SYSTEM, TEMPLATE, and PARAMETER directives introduced.
v0.1.x Series (Initial Release–April 2024)
- v0.1.29: First Windows release
- v0.1.17: Added
ollama --versionflag - v0.1.0: Initial public release. macOS and Linux only. Supported GGUF models via llama.cpp backend.
Breaking Changes by Version
These are the releases where something in your workflow might stop working after an update. I maintain this list because the official release notes sometimes bury breaking changes in minor bullet points.
| Version | What Broke | Migration |
|---|---|---|
| v0.32.0 | Running ollama with no subcommand now starts an interactive agent session | If a script or habit relied on bare ollama, call an explicit subcommand instead (ollama list, ollama run <model>, ollama serve). |
| v0.30.0 | nomic-embed-text now lowercases inputs (to match the model card) | Embeddings shift slightly vs. older Ollama. Re-embed your corpus with v0.30+ so queries and stored vectors match. |
| v0.19 | MLX backend (preview) only activates on Macs with 32GB+ unified memory | None required. On 8GB/16GB Macs Ollama keeps the existing engine; no action needed, you just don't get the MLX speedup yet. |
| v0.5.0 | context field removed from chat API | Use conversation history instead of passing context tokens. See the official docs for the current API. |
| v0.4.0 | Model storage format migration | Automatic on first run. Back up ~/.ollama before updating if you have custom models. |
| v0.3.0 | Modelfile syntax change | Replace FROM base ADAPTER lora.gguf with separate FROM and ADAPTER lines. |
| v0.2.0 | CLI argument changes | --model flag replaced by positional argument. Old: ollama run --model llama2. New: ollama run llama2. |
If you run Ollama behind an application (like Open WebUI or Continue.dev), check that your client version supports the Ollama version you are upgrading to. Open WebUI v0.5+ works with Ollama v0.5+.
Version Comparison Table
This table covers features across major version milestones. Use it to decide what minimum version you need.
| Feature | v0.1 | v0.2 | v0.3 | v0.4 | v0.12–v0.14 | v0.19+ | v0.30+ |
|---|---|---|---|---|---|---|---|
| macOS support | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Linux support | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Windows support | v0.1.29+ | Yes | Yes | Yes | Yes | Yes | Yes |
| Custom Modelfiles | No | Yes | Yes | Yes | Yes | Yes | Yes |
| GPU offloading | No | v0.2.8+ | Yes | Yes | Yes | Yes | Yes |
| OpenAI-compatible API | No | v0.2.5+ | Yes | Yes | Yes | Yes | Yes |
| AMD ROCm support | No | No | v0.3.6+ | Yes | Yes | Yes | Yes |
| Multimodal / vision engine | No | No | No | Basic | Yes (May 2025 engine) | Yes | Yes |
| Web search API | No | No | No | No | Yes | Yes | Yes |
| Cloud / hosted models | No | No | No | No | Yes | Yes | Yes |
| MLX Apple Silicon backend | No | No | No | No | No | Preview (32GB+) | Yes |
How to Roll Back to a Previous Version
Sometimes an update breaks your workflow. Here is how to downgrade safely.
macOS (Homebrew)
# List available versions
brew search ollama
# Install a specific version (example: v0.30.0)
brew install ollama@0.30.0
# If that formula does not exist, install from the GitHub release:
curl -L https://github.com/ollama/ollama/releases/download/v0.30.0/Ollama-darwin.zip \
-o ~/Downloads/Ollama-0.30.0.zip
unzip ~/Downloads/Ollama-0.30.0.zip -d /Applications/
Linux
# Download a specific version binary
curl -L https://github.com/ollama/ollama/releases/download/v0.30.0/ollama-linux-amd64 \
-o /usr/local/bin/ollama
chmod +x /usr/local/bin/ollama
# Restart the service
sudo systemctl restart ollama
# Verify
ollama --version
Docker
# Use a specific tag instead of :latest
docker pull ollama/ollama:0.30.0
docker stop ollama && docker rm ollama
docker run -d --gpus all -v ollama:/root/.ollama \
-p 11434:11434 --name ollama ollama/ollama:0.30.0
Pinning a Version (Preventing Auto-Updates)
# macOS — pin the Homebrew formula
brew pin ollama
# Linux — hold the package if installed via apt
sudo apt-mark hold ollama
# Docker — always use a specific tag, never :latest
Warning: Rolling back from v0.4.0+ to v0.3.x requires restoring ~/.ollama from a backup made before the upgrade, because the storage format migration in v0.4.0 is one-way. Always back up before major version jumps.
Release Cadence and Roadmap
Ollama ships extremely frequently — often several patch releases per week within a minor line, and a new minor version (0.29 → 0.30, etc.) every few weeks. As of August 30, 2026 it sits at v0.33.2 (released August 27), having moved through roughly 33 minor releases since the first public version in 2023.
The clearest recent direction is hardware acceleration and model breadth: the MLX backend for Apple Silicon (v0.19, March 2026) and the ongoing 0.30-series work pairing MLX with an updated llama.cpp engine for wider GPU support, plus a steady stream of new model architectures (Command A, the North family, Cohere2Moe, Gemma 4, and more) added almost every release. With v0.32 the direction widened again: the CLI itself is now an interactive agent, not just a model runner. There is real momentum behind that push — on July 9, 2026 Ollama announced an $88M funding round (Benchmark, Theory Ventures, 8VC, Y Combinator) and reported 8.9M developers using the tool.
Ollama does not publish a fixed public roadmap, so rather than guess at unreleased features, track development directly at github.com/ollama/ollama/releases and the official Ollama blog.
Troubleshooting Update Issues
"ollama: command not found" after update
# macOS — Homebrew may have changed the symlink
brew unlink ollama && brew link ollama
# Linux — verify the binary path
which ollama
# If empty, re-run the install script
curl -fsSL https://ollama.com/install.sh | sh
# Windows — restart your terminal or reboot
# The installer adds Ollama to PATH, but existing terminals
# do not pick it up until reopened
Models disappear after update
Models should survive updates. If they are gone:
# Check if the models directory still exists
ls -la ~/.ollama/models/
# If it is empty, the update may have changed OLLAMA_MODELS path
# Check your environment
echo $OLLAMA_MODELS
# Re-pull missing models
ollama pull llama3.2
Service will not start after update
# Linux — check systemd logs
journalctl -u ollama -n 50
# macOS — check launchd logs
log show --predicate 'process == "ollama"' --last 5m
# Common fix: port conflict from old process
lsof -i :11434
kill -9 <PID>
ollama serve
CUDA / ROCm errors after update
# Verify your GPU driver version
nvidia-smi # NVIDIA
rocm-smi # AMD
# Recent Ollama versions generally need a reasonably current GPU stack:
# NVIDIA: Driver 535+ (CUDA 12.2+)
# AMD: ROCm 6.0+
# (Check the release notes for your exact version — minimums shift over time.)
# If your driver is too old, either:
# 1. Update your GPU driver
# 2. Roll back Ollama to a version that supports your driver
If you hit issues not covered here, check the general Ollama troubleshooting guide or the project's GitHub issues page. For platform-specific installation help, see our Windows installation guide or Mac setup guide.
Staying Current Without Breaking Things
My recommendation: update monthly, not on release day. Let the community shake out bugs for a week or two before you upgrade. Pin your version in production environments and always back up ~/.ollama before major version jumps.
Set a reminder to check your version:
# Add to your .bashrc or .zshrc
alias ollama-check='echo "Installed: $(ollama --version)" && echo "Latest: check https://github.com/ollama/ollama/releases"'
Ollama's release velocity is a strength — they ship features fast and respond to bugs quickly. But that pace means every update deserves a quick test before you trust it with production workloads.
Running Ollama for the first time? Start with the complete Ollama guide for setup instructions, or compare the best Ollama models for your hardware.
Ollama’s running. Here’s what to build with it.
Go from “ollama run” to RAG apps, agents, and fine-tuned models — structured and hands-on. First chapter free.
Stop piecing Ollama together from blog posts
Ollama Mastery is 15 chapters end to end — install, model choice, Modelfiles, GPU offload, the API, and the 20 errors that actually happen. Plus 24 more courses.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARBest Ollama Models 2026: 15 Ranked (Coding, Reasoning, Chat)
- AI on Steam Deck: Run Local LLMs with Ollama on SteamOS
- Air-Gapped AI Deployment: Install Ollama With No Internet
- Best Free Local AI Models to Run With Ollama (No API Key)
- Best Ollama Embedding Models Compared for Local RAG
- Best Ollama Models for 8GB RAM 2026: 12 Tested Local Picks
- Best Ollama Models for AI Agents 2026: Ranked by Tool Use
- Best Uncensored Local LLMs: Abliterated Ollama Models
- Browser-Use + Ollama: A Local Web-Browsing Agent
- Build a Local AI Slack & Discord Bot with Ollama + Python
Comments (0)
No comments yet. Be the first to share your thoughts!