★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
AI Models

Best Claude Model for Coding: Opus vs Sonnet vs Haiku

April 10, 2026
21 min read
Local AI Master Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Picked your coding model? Build a real AI dev workflow. From local copilots to agents that ship code — the structured path, running on your hardware. First chapter free.

Start free
Or own it for life — Lifetime $149, pay once

Published on April 10, 2026 -- 21 min read

TL;DR (July 2026): The best Claude model for most coding is Claude Sonnet 5 -- it is the new default model in Claude Code, priced at $2 input / $10 output per million tokens (intro pricing through August 31, 2026, then $3/$15). It even beats Opus 4.8 on Terminal-Bench 2.1 (80.4% vs 74.6% -- the first Sonnet to beat its Opus sibling on a major coding benchmark). Step up to Claude Opus 4.8 ($5/$25) for the hardest everyday engineering -- it still leads Sonnet 5 on SWE-bench Pro (69.2% vs 63.2%) -- and to Claude Fable 5 ($10/$50, the new tier above Opus) only for long-horizon, multi-file agentic work. Claude Haiku 4.5 ($1/$5) remains the pick for high-volume autocomplete, docstrings, and bulk edits.

The lineup changed in June 2026: Anthropic shipped Claude Fable 5 (June 9), the first model of a new Mythos-class tier that sits above Opus, and Claude Sonnet 5, which replaced Opus as the Claude Code default. So the ladder is now four rungs -- Haiku 4.5, Sonnet 5, Opus 4.8, Fable 5 -- and each one handles code differently. For most developers, Sonnet 5 is what you should reach for first.

One June 2026 event worth knowing about before you build a workflow on any cloud model: US export controls forced Anthropic to suspend Fable 5 access for 19 days (June 12 to July 1) for every user, worldwide. It is back with tightened safety classifiers -- but a coding setup that depends on a single cloud model now has a demonstrated single point of failure that has nothing to do with your uptime. It is one more reason the hybrid local-plus-cloud pattern at the end of this guide keeps gaining ground.

I have been using Claude models for code review, debugging, refactoring, documentation, and project scaffolding. This guide breaks down where each tier shines, where it falls short, and how to choose the right one without treating a benchmark table as permanent truth.

The short answer: use Sonnet for most coding work, Opus for hard reasoning and code review, and Haiku for high-volume simple tasks. Here is why.


The Claude Model Lineup

Anthropic now offers four model tiers through the Claude API and claude.ai:

Claude Fable 5 (new -- above Opus)

The new ceiling, released June 9, 2026. Fable 5 is the first Mythos-class model -- Anthropic says its capabilities "exceed those of any model we've ever made generally available," with the largest leads over Opus 4.8 showing up on longer, more complex tasks. (Its sibling Claude Mythos 5 is the same model without general-use safety measures, restricted to approved organizations.)

Key specs:

  • $10 input / $50 output per million tokens -- exactly 2x Opus 4.8
  • 1M-token context, up to 128K output tokens
  • Adaptive thinking always on; strongest long-horizon agentic performance
  • Anthropic-cited example: Stripe reported a 50-million-line migration compressed from ~two months to a day
  • Overkill (and over-budget) for everyday coding

Claude Opus 4.8

The everyday heavy-lifting tier. Still the model to beat on SWE-bench Pro (69.2%, ahead of Sonnet 5), and half Fable 5's price.

Key specs:

  • $5 input / $25 output per million tokens
  • 1M-token context
  • Strong fit for difficult multi-file reasoning, code review, architecture sessions
  • The escalation tier that most hard problems actually need

Claude Sonnet 5

The workhorse -- and since July, the default model in Claude Code. First Sonnet to beat its Opus sibling on a major coding benchmark (Terminal-Bench 2.1: 80.4% vs 74.6%). This is what most developers should use day-to-day.

Key specs:

  • $2 input / $10 output per million tokens intro pricing through August 31, 2026 (then $3/$15)
  • 1M-token context
  • Faster and 2.5x cheaper than Opus for typical coding responses
  • Best cost-to-quality ratio in the lineup

Claude Haiku 4.5

The speedster. Designed for high-volume, low-latency tasks where speed and cost matter more than maximum quality.

Key specs:

  • 200K-token context
  • $1 input / $5 output per million tokens
  • Fastest Claude tier for simple tasks
  • Excellent for autocomplete, simple code generation, documentation

Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

Coding Benchmarks and How to Use Them

SWE-Bench tests whether a model can resolve real GitHub issues from open-source repositories. The July 2026 numbers tell an unusual story: the tiers now split benchmarks instead of one model sweeping them. Sonnet 5 beats Opus 4.8 on Terminal-Bench 2.1 (80.4% vs 74.6%) while Opus 4.8 stays ahead on SWE-bench Pro (69.2% vs Sonnet 5's 63.2%). Fable 5 tops both -- third-party trackers report ~80% on SWE-bench Pro, though Anthropic hasn't published that exact figure -- and scored highest among frontier models on Cognition's FrontierCode evaluation. Benchmarks are still a snapshot: leaderboards update quickly, and scores change faster than evergreen articles.

For context on how SWE-Bench works, see our SWE-Bench explained guide. Use benchmark tables as a direction signal, not as the only buying decision.

TierBest UseWhy It Wins
Fable 5Longest, hardest multi-file agentic runsLargest leads precisely on long-horizon tasks -- but 2x Opus pricing
Opus 4.8Hard debugging, code review, architectureStill the SWE-bench Pro leader among everyday tiers
Sonnet 5Daily coding, refactoring, testsTerminal-Bench winner at the lowest cost-per-quality in the lineup
Haiku 4.5Bulk edits, docstrings, simple transformationsFastest and cheapest for low-risk work

What this means in practice:

  1. Do not use Opus for every autocomplete-style request. It is usually overkill.

  2. Do not use Haiku for security review or architecture. The task is too high stakes.

  3. Sonnet is the correct default until it fails. Escalate to Opus only when needed.

  4. Re-check official model docs before building cost calculators or sales copy around exact model names and prices.


Pricing Breakdown

Cost Per Task (Typical Developer Usage)

Instead of hard-coding a stale calculator into this page, use the official Anthropic pricing page before estimating cost for a team or production workflow. The ranking is stable even when exact prices change:

TierRelative CostBest Cost Use
HaikuLowestBulk documentation, classification, simple transforms
SonnetMiddleDaily coding, tests, refactors, debugging
OpusHighestCode review, architecture, long agentic sessions

Typical Daily Cost by Developer Profile

ProfileBest DefaultEscalate To
Casual hobby codingSonnetOpus for hard bugs
Active developerSonnetOpus for review and architecture
Bulk automationHaikuSonnet when quality drops
Team workflowSonnet with budget alertsOpus for high-value reviews

The practical math: start with Sonnet, monitor actual token usage, then reserve Opus for tasks where a better answer is worth the extra cost.

Claude Pro Subscription vs API

Claude subscription plans are usually simpler for individual developers because they include the claude.ai interface, project folders, and artifacts. API access is better when you need automation, team attribution, or budget controls.

For teams or heavy API users, direct API access through the Anthropic dashboard gives more control over costs and enables programmatic integration.


Speed vs Quality Tradeoff

Speed matters for coding. Waiting for a response breaks flow, but exact latency depends on context length, output length, region, and current model load.

Response Time by Model (Typical Coding Query: 2K input, 500 output tokens)

TierRelative LatencyBest Speed Use
HaikuFastestInline suggestions and short transforms
SonnetFastInteractive coding and tests
OpusSlowestHard review, architecture, and debugging

Sonnet is fast enough for normal coding flow. It is usually the best balance when you want a thoughtful answer without turning every request into a long agentic session.

Opus has a more noticeable delay, especially with extended reasoning. That delay can be worth it for complex problems, but it interrupts the rapid iteration cycle of active coding.

Haiku feels fastest. For autocomplete-style suggestions and quick lookups, this speed advantage matters.

When Speed Beats Quality

  • Inline code suggestions (Haiku)
  • Quick "what does this function do?" queries (Sonnet)
  • Generating boilerplate (Haiku or Sonnet)
  • Iterating on a prompt -- running 5 versions to find the best one (Sonnet)

When Quality Beats Speed

  • Reviewing a PR for security vulnerabilities (Opus)
  • Debugging a concurrency issue (Opus)
  • Designing an API schema (Opus)
  • Refactoring a 2,000-line module (Opus)

Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

Best Model by Task

Code Generation

Winner: Sonnet

For generating new functions, classes, and modules, Sonnet usually produces clean, well-structured code fast enough for active development. It writes idiomatic code with proper error handling, type annotations, and docstrings without being asked.

Opus generates better code for complex algorithms and edge-case-heavy implementations, but the difference is usually small enough that Sonnet's speed advantage wins for daily use.

Code Review

Winner: Opus

This is where Opus earns its price tag. Given a substantial diff, Opus is the better choice for finding:

  • Logic errors that Sonnet misses
  • Race conditions in concurrent code
  • Security issues (SQL injection, XSS, path traversal)
  • Performance problems (N+1 queries, unnecessary allocations)
  • Architectural concerns (coupling, SRP violations)

Sonnet catches obvious issues well, but Opus is the better default when missed subtle bugs can cost hours or days.

Debugging

Winner: Opus

Debugging requires understanding state, control flow, and the interaction between components. Opus excels here because extended reasoning helps it trace through execution paths systematically. Feed it a stack trace, the relevant source files, and a description of expected vs actual behavior, and it often narrows the root cause faster.

Sonnet is adequate for straightforward bugs, but Opus is a better fit for concurrency bugs, memory leaks, and issues that span multiple modules.

Refactoring

Winner: Opus for large refactors, Sonnet for small ones

Renaming a variable, extracting a method, simplifying a conditional -- Sonnet handles these well. For refactoring an entire module, splitting a monolith, or migrating a codebase to a new pattern, Opus produces better results because it maintains awareness of how changes ripple through the codebase.

Test Writing

Winner: Sonnet

Test generation is relatively formulaic: read the function signature, understand the edge cases, write assertions. Sonnet does this well and fast. Opus writes more thorough tests for complex integration scenarios, but the difference rarely justifies using the premium tier for every test.

Documentation

Winner: Haiku

Writing docstrings, README updates, API documentation, and inline comments is exactly the kind of task where Haiku's cost and speed advantages shine. The quality is good enough for documentation, and you can run it across an entire codebase for pennies.


Claude Code CLI

Claude Code is Anthropic's official command-line tool for agentic coding. It connects Claude directly to your terminal and filesystem.

What Claude Code Does

  • Reads and writes files across your project
  • Runs terminal commands (build, test, lint)
  • Creates and manages git commits
  • Searches codebases with grep/glob
  • Handles multi-step refactoring autonomously

Default Model and Overrides

Since July 2026, Claude Code defaults to Sonnet 5 on Pro and Team plans -- a meaningful change, because the previous default was the Opus tier. In practice the default is now right for most sessions, and you escalate to Opus 4.8 (or Fable 5, if your plan includes it) for the hardest agentic runs. Each tool call (read file, write file, run command) costs tokens, so agentic sessions can get expensive on large projects.

# Install Claude Code
npm install -g @anthropic-ai/claude-code

# Start an agentic coding session
claude "Refactor the auth module to use JWT instead of sessions"

# Override to Sonnet for simpler tasks
claude --model sonnet-current "Add input validation to the user form"

Claude Code vs Cursor vs Copilot

FeatureClaude CodeCursorGitHub Copilot
InterfaceTerminal/CLIVS Code forkIDE extension
Typical model choiceOpus / SonnetSonnet / GPT-class modelCopilot model family
Agentic capabilityFull (files, terminal, git)Moderate (file edits)Limited (inline suggestions)
Multi-file editingNativeComposer modeLimited
CostPay per useSubscriptionSubscription
OfflineNoNoNo
Best forComplex refactoring, debuggingDaily coding, all tasksInline completion

For a deeper comparison of coding tools, see our AI coding tools comparison.


API Integration for Developers

If you are building coding tools or integrating Claude into your development workflow:

Basic API Call for Code Generation

import anthropic

client = anthropic.Anthropic()

message = client.messages.create(
    model="sonnet-current",
    max_tokens=4096,
    messages=[
        {
            "role": "user",
            "content": """Review this Python function for bugs and suggest improvements:

def calculate_discount(price, discount_percent):
    if discount_percent > 100:
        return 0
    final_price = price - (price * discount_percent / 100)
    return final_price"""
        }
    ]
)

print(message.content[0].text)

Choosing Model by Task in Code

def get_model_for_task(task_type: str) -> str:
    """Select a Claude model tier based on task complexity.

    Replace these aliases with current model IDs from Anthropic docs.
    """
    model_map = {
        "autocomplete": "haiku-current",
        "generate": "sonnet-current",
        "review": "opus-current",
        "debug": "opus-current",
        "refactor_small": "sonnet-current",
        "refactor_large": "opus-current",
        "test": "sonnet-current",
        "document": "haiku-current",
    }
    return model_map.get(task_type, "sonnet-current")

Extended Thinking for Complex Problems

# Opus with extended thinking for debugging
message = client.messages.create(
    model="opus-current",
    max_tokens=16000,
    thinking={
        "type": "enabled",
        "budget_tokens": 10000  # Allow up to 10K tokens of reasoning
    },
    messages=[
        {
            "role": "user",
            "content": "This async Python function deadlocks intermittently. Analyze the code and identify all potential race conditions: [code here]"
        }
    ]
)

Claude vs Other Coding Models

Claude is not the only serious coding option. OpenAI, Google, and local open-weight models are all competitive depending on your workflow. Exact benchmark rankings change often, so the safer comparison is by workflow fit.

Strengths by Provider

Claude (Anthropic):

  • Strong multi-file understanding and refactoring
  • Strong fit for code review and agentic coding
  • Claude Code provides a terminal-native coding workflow

OpenAI models:

  • Strong ecosystem integration across developer tools
  • Strong general coding and structured-output workflows
  • Good choice if your stack already uses OpenAI APIs

Gemini models:

  • Strong long-context and multimodal workflows
  • Useful when screenshots, UI diagrams, or very large context windows matter
  • Strong fit for Google Cloud-heavy teams

Which to Choose?

For code review and agentic debugging: Claude Opus. For daily cloud coding: Claude Sonnet. For ecosystem integration: choose the model family already wired into your tooling. For local/private coding: None of these -- see our best local AI coding models instead, or the section below.


When a Local Model Wins (and Saves Money)

This is a guide to the best Claude model for coding, and for the hardest agentic and review work Claude is genuinely worth it. But it would be dishonest to pretend a cloud API is always the right answer. A lot of everyday coding now runs perfectly well on an open-weight model on your own machine -- no per-token bill, no data leaving your laptop, and no internet required. By late 2026 that bar has moved a long way.

Choose a local model instead of Claude when one of these is true:

  • Privacy / data sovereignty. Proprietary code, client work under NDA, healthcare or regulated data -- if it cannot leave your network, a local model is not a compromise, it is the requirement. Your code never touches a third-party server.
  • Cost at volume. Claude Sonnet 5 is $2 input / $10 output per million tokens on intro pricing ($3/$15 after August 31), Opus 4.8 is $5/$25, and Fable 5 is $10/$50. That is cheap for occasional use, but if you run autocomplete and bulk edits all day, the meter adds up fast. A local model is free per token after the hardware you already own.
  • Availability risk. Fable 5 was suspended for every user worldwide for 19 days in June 2026 under a US export-control order. Local weights on your own disk cannot be switched off by anyone.
  • Offline / air-gapped. On a plane, on a locked-down corporate network, or in a secure facility, a cloud API simply is not available. A local model keeps working.
  • High-volume low-stakes work. Autocomplete, docstrings, boilerplate, simple transforms -- the exact tasks where you would otherwise reach for Haiku -- run fine locally and cost nothing per call.

The local coding models that are now good enough:

  • Qwen3-Coder -- the current open-weight standout for agentic coding. The 30B-class variant runs well on a 24GB GPU or a 32GB+ Mac and holds up across refactors and multi-file edits.
  • Devstral Small 2 (24B) -- Mistral's purpose-built coding-agent model. It scores about 68% on SWE-bench Verified -- the strongest open-weight model in its size class, beating many 70B-class competitors -- and Mistral designed it to run on a single RTX 4090 (24GB) or a 32GB Mac. See our Devstral deep dive for setup and benchmarks.
  • Codestral -- Mistral's fill-in-the-middle specialist, built for fast inline completion with a long context window. A strong pick for the autocomplete workload where Haiku would otherwise be your cloud default.

If you have a 12-16GB GPU, the 14B-class tier is the sweet spot -- big enough to be genuinely useful for refactors and tests, small enough to run on a mid-range card. Our roundup of the best 14B coding models ranks the current options. For the full picture across every VRAM tier, see the best local AI models for programming pillar.

Not sure which local model fits your hardware and task? Use our coding model router: pick what you are coding (autocomplete, refactor, debug, agentic, tests) and your GPU, and it returns the open-weight model that actually fits -- and tells you honestly when the job has outgrown your hardware and Claude is the better call.

How to actually run a local model inside Claude Code

The section above is about whether to go local. This is the how, because it is the thing most people are really asking when they search for the best model for Claude Code.

Ollama added launcher support for coding agents over the summer of 2026. As of Ollama v0.32.11 you can start Claude Code against a local model in one command:

# start Claude Code backed by a local model
ollama launch claude --model qwen2.5-coder:32b

Ollama v0.33.0 went further and added Claude Desktop support, letting you configure the desktop client to use Ollama as a third-party gateway provider — so the same local weights back both the CLI and the desktop app. Both changes are documented in Ollama's release notes.

Two things decide whether this feels good or frustrating:

  • Pick the model by your VRAM, not by the leaderboard. A 24GB card runs qwen2.5-coder:32b; 16GB runs devstral; 12GB runs qwen2.5-coder:14b. Our best Ollama model for coding page has the pick and pull command for each tier.
  • Raise the context length. Agentic coding reads several files at once, and Ollama's default context is conservative. A local model that cannot see enough of your repository will look far worse than it is. Set it per-model rather than globally.

For the full walkthrough including config and the failure modes, see running Claude Code offline with Ollama.

Honest limits -- where Claude still leads

Local is not a clean sweep. The frontier cloud models still win on the hardest ~10-20% of tasks: tangled architectural refactors, subtle concurrency and memory bugs, and long agentic sessions that span many files. If you are on weak hardware (8GB VRAM or CPU-only), you are capped at ~7B models and cloud will usually be the smarter route. The most productive setups are hybrid: a local model for the routine 80% of edits and completions, Claude Sonnet 5 or Opus 4.8 for the gnarly 20% where deeper reasoning earns its price.


Real Code Quality Differences

Abstract benchmarks are useful, but what do the quality differences actually look like? Here are examples from the same prompt given to all three Claude tiers.

Task: "Write a rate limiter middleware for Express.js"

Haiku output (simplified): Generates a basic token bucket implementation. Works but uses a simple in-memory object, does not handle distributed environments, and misses edge cases like clock drift. About 30 lines.

Sonnet output: Produces a sliding window rate limiter with configurable limits per route. Includes proper error responses (429 with Retry-After header), optional Redis backend for distributed deployments, and TypeScript types. About 80 lines with comments.

Opus output: Everything Sonnet produces, plus: race condition handling in the Redis backend, graceful degradation when Redis is unavailable, separate limits for authenticated vs anonymous users, IP-based and user-based limiting, and a test suite. About 150 lines with thorough documentation.

The pattern repeats across tasks: Haiku gives you the minimum viable implementation. Sonnet gives you a strong first draft. Opus gives you the version that spends more attention on failure modes.


When to Use Each Model

Use Claude Haiku When:

  • Running autocomplete / inline suggestions at scale
  • Generating documentation across a large codebase
  • Processing many small, independent code tasks in parallel
  • Budget is the primary constraint
  • The code is simple enough that quality differences are minimal
  • You need sub-second response times

Use Claude Sonnet When:

  • Writing new features during active development
  • Generating unit tests and integration tests
  • Small to medium refactoring (single file, few files)
  • Interactive debugging with rapid iteration
  • You want the best quality-per-dollar ratio
  • Daily coding across all task types

Use Claude Opus When:

  • Reviewing pull requests for production code
  • Debugging complex, multi-component issues
  • Large-scale refactoring (module restructuring, migration)
  • Architectural design and system design discussions
  • Security audits and vulnerability analysis
  • Agentic tasks via Claude Code

The Practical Workflow

Most experienced developers use all three tiers throughout their day:

  1. Morning code review: Opus reviews overnight PRs
  2. Active development: Sonnet for code generation and quick debugging
  3. Test writing: Sonnet generates test suites
  4. Documentation: Haiku generates docstrings in bulk
  5. End-of-day refactor: Opus handles the complex cleanup

This mixed approach keeps daily costs in the $5-15 range while getting maximum quality where it matters most.


Conclusion

The best Claude model for coding is not a single model -- it is knowing when to use each tier. Sonnet covers most daily coding at a lower cost than Opus. Opus is worth the premium for tasks where subtle quality differences have outsized impact. Haiku earns its place for high-volume, low-complexity work.

If you are forced to pick just one: Claude Sonnet 5. It is the Claude Code default for a reason -- the most practical choice for developers who need a capable AI partner throughout the workday, at the lowest cost per unit of quality in the lineup.

If money is no object and you want the highest reasoning tier: Claude Opus 4.8 for review, debugging, architecture, and long agentic coding sessions -- and Claude Fable 5 only when a task is genuinely long-horizon enough to justify twice the price.


Want to compare Claude with local alternatives that run on your own hardware? Check our best AI coding models ranking or set up a free local AI coding assistant with Continue.dev.

🎯
AI Learning Path

Picked your coding model? Build a real AI dev workflow.

From local copilots to agents that ship code — the structured path, running on your hardware. First chapter free.

Or own it for life — Lifetime $149 $599, pay once

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsClaudeOpusSonnetHaikuCodingAI ProgrammingAnthropic

Local AI Master Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on AI Models for Coding
See the full Best Local AI for Coding guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

📅 Published: April 10, 2026🔄 Last Updated: July 20, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Was this helpful?

📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Picked your coding model? Build a real AI dev workflow.

From local copilots to agents that ship code — the structured path, running on your hardware. First chapter free.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators