G
GEO Toolbox
claude-sonnet-5claudeanthropicsonnet-5llmguide

Claude Sonnet 5: Pricing, Benchmarks & Real-World Review (2026)

Claude Sonnet 5 is out. Pricing, benchmarks vs Sonnet 4.6 and Opus 4.8, the real per-task cost at high effort, and what the first independent reviews reveal.

Samy Ben SadokSamy Ben Sadok17 min read
In this post13 sections

Claude Sonnet 5 launched June 30, 2026. Months of leaked rumors under the wrong codename preceded it, and the real model lands close to what those rumors promised: most of Opus 4.8's agentic ability, at a meaningfully lower price, with a few real caveats the launch posts gloss over.

What Is Claude Sonnet 5?

Claude Sonnet 5 is Anthropic's mid-tier Claude model, released June 30, 2026. It sits between the fast, low-cost Haiku 4.5 and the higher-tier Opus and Fable 5 models, and Anthropic calls it "the most agentic Sonnet model yet."

Note that the Opus tier has moved on since this article was written. Anthropic released Claude Opus 5 on July 24, 2026, replacing Opus 4.8 at the same $5 and $25 per million tokens. The Opus 4.8 comparisons below still describe the model Sonnet 5 launched against.

That framing is specific, not marketing filler. Sonnet 5 can build a multi-step plan, decide which tools it needs (a browser, a terminal, a file editor), execute that plan with minimal hand-holding, and check its own output before handing back a result. Anthropic's own write-up describes testers asking it to investigate a bug. Unprompted, it wrote a reproducing test, implemented the fix, then stashed the change to confirm the bug actually came back without it, all in one pass.

Sonnet 5 closes most of the gap to Opus 4.8 while staying at Sonnet-tier pricing. For many agentic and coding tasks, Claude users now have a meaningfully cheaper option that doesn't feel like a downgrade. That doesn't hold at every effort setting, though, and the real cost math gets its own section below.

Bar chart of SWE-bench Pro scores: Sonnet 4.6 at 58.1%, Sonnet 5 at 63.2%, and Opus 4.8 at 69.2%
On SWE-bench Pro, Sonnet 5 clears Sonnet 4.6 by five points but still trails Opus 4.8 by six. It closes most of the gap, not all of it.

No, You're Not Misremembering "Sonnet 5" from Months Ago

If "Claude Sonnet 5" sounds familiar, you're thinking of "Fennec." In early February 2026, a Google Vertex AI error log exposed a model identifier, claude-sonnet-5@20260203, alongside that internal codename. It looked like proof a launch was imminent, and a wave of speculative coverage ran with it for months, including fabricated benchmark tables and at least one April Fool's Day satire post with invented scores like 92.4% on SWE-bench Verified. None of it came from Anthropic.

What actually happened: that leaked checkpoint shipped as Claude Sonnet 4.6 on February 17, 2026, not as Sonnet 5. Sonnet 5 itself launched four months later, on June 30, 2026, with an official announcement, a system card, and the numbers below. If you've been holding onto a "Fennec" spec sheet, throw it out, it was describing a different model.

Claude Sonnet 5 Pricing

Claude Sonnet 5 launched with introductory pricing that Anthropic made permanent in August 2026, cancelling the planned increase:

PeriodInput (per million tokens)Output (per million tokens)
Introductory (announced through Aug 31, 2026)$2.00$10.00
Now permanent (increase to $3/$15 cancelled)$2.00$10.00

For comparison, Claude Opus 4.8 runs $5 per million input tokens and $25 per million output tokens (see our full Claude pricing guide). Sonnet 5's price is roughly 60% cheaper than Opus 4.8 on both ends.

The tokenizer changed too, so the price-per-token comparison overstates the savings. Sonnet 5 runs on an updated tokenizer, the same change Anthropic introduced with Opus 4.7, which processes the same input text into roughly 1.0 to 1.35 times as many tokens depending on content type. Anthropic priced the launch rate specifically so the transition lands as "roughly cost-neutral" despite that change. The practical version of that math, what it means for your actual workload, gets its own section further down.

Anthropic also says it has raised rate limits across Chat, Cowork, Claude Code, and the Claude Platform to accommodate the higher token usage that comes with running Sonnet 5 at higher effort levels (if you code with it, our guide to cutting Claude Code token costs covers how to keep that in check) (a related platform-wide rate-limit restructuring took effect April 26, 2026, ahead of this launch).

Claude Sonnet 5 Benchmarks: vs Sonnet 4.6 and vs Opus 4.8

Anthropic published five benchmark comparisons at launch. Sonnet 5 beats its predecessor on every single one, and closes most, though not all, of the distance to Opus 4.8.

BenchmarkSonnet 4.6Sonnet 5Opus 4.8What it measures
SWE-bench Pro58.1%63.2%69.2%Long-horizon agentic software engineering
Terminal-Bench 2.167.0%80.4%74.6%Command-line tool use
Humanity's Last Exam (with tools)46.8%57.4%57.9%Graduate-level multidisciplinary reasoning
OSWorld-Verified78.5%81.2%83.4%Computer use (operating-system tasks)
GDPval-AA v2-1,618 pts1,615 ptsReal-world professional knowledge work

SWE-bench Pro is the hardest, most coding-specific test of the five, and it's the clearest Opus win: Opus 4.8 leads 69.2% to 63.2%, and Anthropic is upfront that Opus remains "the model of choice for higher accuracy" on this kind of work. Opus also leads on OSWorld-Verified and edges Humanity's Last Exam. But it doesn't sweep the table, and the popular "Opus wins everything" summary is wrong on two rows. On Terminal-Bench 2.1, Anthropic's own Opus 4.8 launch post puts it at 74.6% — behind Sonnet 5's 80.4%, making Sonnet 5 the stronger Claude model on command-line tool use. And on GDPval-AA v2, which scores real-world knowledge tasks rather than coding puzzles, Sonnet 5 numerically edges Opus 4.8, 1,618 to 1,615, a gap close enough to read as a tie. So the honest read is narrower than the headline: Opus leads pure coding, Sonnet 5 wins terminal work outright, and the two trade the rest.

Anthropic also updated its grading for Humanity's Last Exam and OSWorld-Verified at this launch, and retroactively re-scored Sonnet 4.6 under the new methodology: its HLE score moved to 46.8% with tools, its OSWorld score to 78.5%. That's why the Sonnet 4.6 numbers above may not match what you remember from Sonnet 4.6's own launch post. This table uses the regraded baseline, the fairer comparison.

Third-party testing points the same direction. Cursor ran its own internal benchmark, CursorBench, and reported Sonnet 5 scoring 57% against Sonnet 4.6's 49%. That's Cursor's proprietary test, not one of Anthropic's published evaluations, but the direction matches everything else here.

New Capabilities: Effort Levels and Agentic Behavior

Sonnet 5 runs with adaptive thinking on by default. You choose how hard it thinks through an effort setting, ascending from Low to Medium to High to Extra High to Max. Extra High is new for the Sonnet line: Sonnet 5 is the first Sonnet-class model to offer it, though Max, the actual ceiling, was already available on Sonnet 4.6. Effort defaults to High on both the API and in Claude Code. Higher effort means more reasoning, more tool calls, and, as covered above, more spend.

The model carries a 1 million token context window and a maximum output of 128,000 tokens per response, extendable to 300,000 via a batch-API beta header for long-running jobs. Training data runs through January 2026.

In practice, according to Anthropic's named early-access partners, that translates to fewer stalled multi-step jobs. Zapier handed Sonnet 5 a two-part task: update Salesforce account tiers, then send a launch announcement to enterprise contacts. It finished the whole thing end to end, something that used to stall halfway with prior models. Legal-tech platform Eve said Sonnet 5 "sits on the Pareto frontier" for its plaintiff-law tasks, citing a price-to-performance ratio that made switching straightforward. ClickHouse said the model reasons in tighter steps and gets users to answers noticeably faster when exploring live data.

Anthropic also points to strength on "brownfield" code: the messy, already-shipped parts of a codebase nobody wants to touch, like race conditions and hidden test failures. The claim is that it traces a failure back to its actual root cause instead of patching the visible symptom. That's a different skill from generating new code from scratch, and it's a useful one for agentic AI work specifically, because a fix either holds under the original failing test or it doesn't.

One spec worth a caveat: despite the 1 million token window, some early testers report the model losing track of details across very long documents or codebases, more than they expected given the headline number. Treat the context window as a ceiling on how much you can feed the model, not a guarantee it will weigh every part of that input equally well.

What the First Independent Tests Found

Anthropic's own benchmarks are one thing; the first outside tests, run in the days after launch, add nuance the launch posts skip. The clearest example is CodeRabbit's automated code-review benchmark, worth reading closely because it doesn't split cleanly for or against the model.

On the plus side, Sonnet 5's review precision jumped to 38-40%, up from about 29% on Sonnet 4.6: its findings were more often real bugs than noise, and its comments were cleaner. The catch is a genuine surprise. On raw bug-catching, Sonnet 5 actually flagged fewer real bugs than its predecessor, roughly 50% of the seeded set versus 63% for Sonnet 4.6. Cleaner output, fewer false alarms, but lower recall on the defects that mattered. If your job is finding bugs rather than writing code, that trade-off is worth knowing before you switch. It's also a concrete answer to the naming critics: "the most agentic Sonnet yet" is not the same claim as "better than 4.6 at everything."

Two more practical notes from the same test. At maximum effort, Sonnet 5 posted three to four times as many nitpicks and roughly doubled the cost without finding meaningfully more bugs, so CodeRabbit recommends running it at medium effort to capture most of the benefit without flagship-level spend, the same conclusion the cost math below points to. And because of its extended thinking, it runs slower than Sonnet 4.6 and occasionally rewrites its own plan mid-task, which makes it a poor fit for high-volume pipelines on a tight latency budget.

Safety and Cybersecurity Changes

Anthropic's pre-deployment safety evaluations found Sonnet 5 an overall improvement on Sonnet 4.6. It's better at refusing malicious requests, more resistant to prompt-injection hijack attempts, and lower on hallucination and sycophancy.

On the automated behavioral audit that tests for misuse cooperation and deception, Sonnet 5 scored safer than Sonnet 4.6 overall. Anthropic notes it showed somewhat higher rates of misaligned behavior on that specific assessment than the more capable Opus 4.8, a reminder that "safer than its predecessor" and "as safe as the flagship" are different claims.

Cybersecurity capability is where Anthropic drew a deliberate line. Sonnet 5 wasn't trained on cybersecurity tasks, and on evaluations that test for genuinely dangerous skills, like developing working software exploits, it performs substantially worse than Opus 4.8 and Mythos 5. In one test built around real Firefox vulnerabilities, Sonnet 5 never produced a single full working exploit, though it did show a slightly higher rate of partial success than Sonnet 4.6. Anthropic attributes that to general intelligence gains rather than any cyber-specific training.

Because of that small uptick, Sonnet 5 ships with real-time cyber safeguards on by default, matching the protections already running on Opus 4.7 and 4.8. They're lighter than the restrictions Anthropic put on Fable 5, which block a much wider range of cybersecurity-adjacent tasks. Fable 5 and its research sibling Mythos 5 were also briefly pulled under a US government export restriction over cybersecurity risk in June 2026, but that order was lifted on July 1 and Fable 5 is generally available again (Mythos 5 stays gated to approved organizations). Sonnet 5 launched without that baggage, positioned as the safer, broadly deployable option. Organizations that need reduced guardrails for legitimate security work can apply through Anthropic's Cyber Verification Program, available today on the native Claude Platform, AWS, and Microsoft Foundry, with Google Vertex support coming soon.

The Real Cost of Switching to Sonnet 5

Headline pricing puts Sonnet 5 at 40-60% cheaper than Opus 4.8. Whether your actual bill drops that much depends on two things the price-per-token comparison doesn't capture.

The first is the tokenizer change from the pricing section above: more tokens for the same input, by design, depending on content type. Anthropic built that into the launch pricing to land "roughly cost-neutral," but cost-neutral on average isn't the same as cost-neutral for your specific workload. Code-heavy or non-English content tends to sit at the higher end of that multiplier.

The second factor is bigger, and it's specifically a Sonnet-5-vs-Sonnet-4.6 issue, not a Sonnet-5-vs-Opus-4.8 one. The Decoder's analysis of the launch makes the point plainly: because Sonnet 5 works more agentically, it's likely to chew through more tokens per task than its predecessor, so even at an unchanged per-token rate, running it could end up costing more than Sonnet 4.6 did for the same job. Anthropic saw the same pattern when Opus moved from 4.6 to 4.7. At higher effort settings specifically, that extra token volume compounds, so the per-task savings against any prior model, Sonnet or Opus, can shrink well below what the rate card alone suggests.

There's a tokenizer trap layered on top. One post-launch cost analysis ran a fixed daily workload through the numbers: on the $2/$10 rate with moderate tokenizer inflation, it comes out around 20% cheaper than the same job on Sonnet 4.6. That analysis also modelled a scheduled September 1 rise to $3/$15 that, combined with high-end tokenizer overhead, would have pushed the identical workload 20-35% above the old Sonnet 4.6 baseline, but Anthropic cancelled that increase in August 2026 and kept $2/$10 as the permanent rate, so the crossover it warned about no longer happens. What is still worth watching is the tokenizer overhead itself: budget against your actual token volume at high effort, not just the per-token rate.

We saw the token-volume side first-hand. Running our own multi-phase, multi-agent content workflow on Sonnet 5 the day it launched, the session needed roughly twice the token budget that same workflow usually takes on Opus, enough to trigger a mid-run context-window compaction we hadn't hit running the identical process before. One internal data point, not a benchmark, but it lines up with the pattern Anthropic, The Decoder, and CodeRabbit's max-effort test all describe.

At Low or Medium effort, the savings versus Opus 4.8 hold up clearly, and Medium is where independent testing lands as the sweet spot. At High effort and above, compare actual per-task spend, not just the price-per-token, before assuming you've cut your bill.

Where You Can Use Claude Sonnet 5 Today

Sonnet 5 is the new default model for Free and Pro plan users, and it's available to Max, Team, and Enterprise users as well. Developers can reach it through the Claude API as claude-sonnet-5, and it's live in Claude Code from day one. On the infrastructure side, it's available through the Claude Platform on AWS, Amazon Bedrock, Microsoft Foundry, and Google Vertex AI, with only the Cyber Verification Program itself still rolling out on Vertex.

Third-party tools picked it up the same day. GitHub Copilot made Sonnet 5 generally available at launch across Copilot Pro, Pro+, Max, Business, and Enterprise tiers, with a gradual rollout across VS Code, Visual Studio, the Copilot CLI, JetBrains, and Xcode. Cursor added it the same day too, complete with its own CursorBench numbers.

Early users flagged two rollout snags. While the API and Claude Code had Sonnet 5 live immediately, some reported it wasn't yet selectable in the Claude Desktop app or the Claude Code VS Code extension in the first hours after launch, a sequencing gap rather than a sign the model isn't really live. Separately, Anthropic didn't reset usage limits for the launch, so some users who had already hit their Sonnet 4.6 usage cap couldn't try the new model right away. The core platforms (chat, API, and Claude Code) were confirmed working from the start.

Early Reactions: What Developers Are Saying

Day-one reaction split along predictable lines. Developers running agentic coding workflows were enthusiastic about the jump in tool use and multi-step follow-through, with several early posts calling it the new default choice for Claude Code work that doesn't need Opus-level accuracy. Others pushed back on the value proposition at higher effort settings, for the cost reasons covered above, and a vocal subset argued the jump from Sonnet 4.6 didn't earn the "5" in the name since it doesn't claim a new best-in-class coding score the way some past Sonnet releases did. It's a fair naming argument, not a factual one: the benchmark gains over Sonnet 4.6 are real regardless of what the release was called. It's the same pattern we saw with Grok 5, where a confirmed release still gets picked apart by critics looking for reasons the number shouldn't have gone up.

On how it stacks up against OpenAI and Google: a true apples-to-apples benchmark still doesn't exist, because the three vendors don't share a common test suite and Anthropic's page carries no cross-vendor numbers. But in the days since launch a rough consensus has settled. Sonnet 5 is widely rated the best default pick among broadly available frontier models, almost entirely on price-to-performance. OpenAI's GPT-5.6 Sol posts the highest raw agentic scores anyone has published (its top tier hits about 92% on Terminal-Bench 2.1, above every Claude model), and it stopped being a hypothetical on July 9, 2026, when it went generally available. So the argument for Sonnet 5 is no longer that the alternative is out of reach; it is the bill. Sol runs $5/$30 per million tokens against Sonnet 5's $2/$10 rate, two and a half times as much on input and three times on output, for the top of the leaderboard. If a workload genuinely needs those scores, Sol is now there to be bought. If it doesn't, and for most teams it doesn't, Sonnet 5 gets you most of the way for a fraction of it. Google's Gemini 3.1 Pro stays the pick for very long context and multimodal work. For a direct Claude-vs-ChatGPT breakdown, see our comparison of the two models; we'll fold Sonnet 5 in there as the shared-benchmark picture firms up.

What This Means If You're Tracking AI Visibility

Here's a live example of why this matters. Hours after Sonnet 5's launch, we checked how different AI engines answered questions about it. Most correctly found and cited Anthropic's real announcement. One major engine, however, was still confidently describing Sonnet 5 as unreleased, citing months-old "Fennec" rumor coverage as its source, even with live web search turned on. Search-enabled doesn't automatically mean current, and on a fast-moving release, an AI engine can keep citing stale information well after the facts changed. Re-checking days later, most engines have corrected the launch date, but the old "Fennec" pages, complete with invented benchmark numbers, still circulate in search results and get pulled into some answers, exactly the lag that shifts what buyers read while a topic is still moving.

That's the exact gap brand citation tracking exists to catch, and not just on launch days. A cheaper, more agentic Claude model lowers the cost of running autonomous research, shopping, and support agents at scale, which means more of the buying journey happens inside an AI conversation rather than a search results page. In our experience, the days right after a major model update are when AI engines' answers shift the most, sometimes because a model genuinely reasons differently, sometimes because of exactly the kind of stale-data lag we just saw.

If you want to see whether ChatGPT, Gemini, Claude, Grok, and the other major engines are actually citing your brand, or surfacing your competitors' sources instead, Citation Interceptor tracks that across eight AI engines including Claude, and flags the gaps in what those engines cite when your buyers ask.

Frequently Asked Questions

Is Claude Sonnet 5 available now?

Yes. It launched June 30, 2026, and is live as the default model for Free and Pro plans, available to Max, Team, and Enterprise users, and accessible via the Claude API and Claude Code.

How much does Claude Sonnet 5 cost?

$2 per million input tokens and $10 per million output tokens. Anthropic first announced that as introductory pricing through August 31, 2026, but in August 2026 it cancelled the planned rise to $3/$15 and made $2/$10 the permanent rate.

Is Claude Sonnet 5 better than Opus 4.8?

Not across the board. Opus 4.8 still leads on the hardest agentic coding benchmark, SWE-bench Pro, and is Anthropic's recommended choice for the highest-accuracy work. On knowledge-work tasks (GDPval-AA v2), the two are essentially tied. Which one wins depends on the task and the effort level you're running at.

Is Claude Sonnet 5 better than Sonnet 4.6?

Mostly, but not universally. It wins every published benchmark over 4.6 and is markedly better at multi-step agentic follow-through. The exception showed up in independent testing: on automated code review, Sonnet 5 caught fewer real bugs than 4.6 (about 50% versus 63%) despite cleaner, more precise output, and it uses more tokens per task. For generating and shipping code it's a clear upgrade; for pure bug-hunting on a tight budget, 4.6 can still hold its own.

What effort level should I run Claude Sonnet 5 at?

Effort defaults to High on the API and in Claude Code, but independent testing points to Medium as the sweet spot for most work: it captures most of the quality without the token bill. Maximum effort roughly doubled cost in one code-review benchmark without finding meaningfully more bugs, so reserve High and above for tasks that genuinely need the extra reasoning.

Can I use Claude Sonnet 5 for free?

Yes. It's the default model on Claude's Free plan as of launch, with no separate signup required.

What happened to "Fennec," the leaked Sonnet 5 from earlier this year?

"Fennec" was the internal codename attached to a model identifier that leaked via Google Vertex AI logs in February 2026. That specific checkpoint shipped as Claude Sonnet 4.6, not Sonnet 5, and months of rumor coverage built on top of it, including invented benchmark numbers, never came from Anthropic. The real Claude Sonnet 5 launched June 30, 2026, with its own verified specs.

Is Claude Sonnet 5 available in Claude Code?

Yes, from day one, alongside the Claude API and Claude Platform. Some other surfaces, like the Claude Desktop app, had a brief rollout lag in the hours immediately after launch.

Sources

  • Introducing Claude Sonnet 5 - Anthropic - the official launch announcement: pricing, benchmarks, safety evaluations, and customer quotes - anthropic.com/news/claude-sonnet-5
  • Anthropic launches Claude Sonnet 5 as a cheaper way to run agents - TechCrunch - competitive pricing context and business framing - techcrunch.com/2026/06/30/anthropic-launches-claude-sonnet-5-as-a-cheaper-way-to-run-agents
  • Anthropic's new Claude Sonnet 5 closes the gap to the pricier Opus model series - The Decoder - full benchmark breakdown and the real-world token-cost analysis - the-decoder.com/anthropics-new-claude-sonnet-5-closes-the-gap-to-the-pricier-opus-model-series
  • Claude Sonnet 5 review: Should you switch? - CodeRabbit - independent post-launch code-review benchmark: precision gains, the bug-recall regression vs Sonnet 4.6, and the medium-effort recommendation - coderabbit.ai/blog/claude-sonnet-5-review
  • Claude Sonnet 5 Pricing 2026: The Hidden Costs and Real Savings - Finout - the worked cost example modelling the September 1 standard-rate crossover Anthropic later cancelled - finout.io/blog/claude-sonnet-5-pricing-2026-the-hidden-costs-and-real-savings-behind-the-cost-neutral-launch
  • Anthropic launches Claude Sonnet 5 model on Claude and APIs - TestingCatalog - Cursor integration and the CursorBench third-party benchmark - testingcatalog.com/anthropic-launches-claude-sonnet-5-model-on-claude-and-apis
  • Claude Sonnet 5 is generally available for GitHub Copilot - GitHub Changelog - GitHub Copilot day-one availability and plan tiers - github.blog/changelog/2026-06-30-claude-sonnet-5-is-generally-available-for-github-copilot
  • Claude Sonnet 5 "Fennec" leak: what the Vertex AI logs actually show - Dev Community - the leaked Vertex AI log and the checkpoint identifier - dev.to/marc0dev/claude-sonnet-5-fennec-leak-what-the-vertex-ai-logs-actually-show-3ho5
  • Claude Sonnet 5 "Fennec" Leak: What Actually Launched as Claude Sonnet 4.6 - NxCode - confirms the leaked checkpoint shipped as Sonnet 4.6 on February 17, 2026 - nxcode.io/resources/news/claude-sonnet-5-fennec-leak-2026
  • Anthropic just dropped Claude Sonnet 5 - Dev Community (April Fool's satire, self-labeled) - the fictional benchmark post referenced as an example of the rumor cycle's reach - dev.to/best_codes/anthropic-just-dropped-claude-sonnet-5-and-the-benchmarks-are-kind-of-insane-3ppc

Get GEO insights in your inbox

One email when we publish something worth reading.

Keep reading