All posts
12 min read

Claude Sonnet 5.5 Review: Same Price, Far Less Waste

Claude Sonnet 5.5 keeps Sonnet 5's $2/$10 pricing, costs far less than Opus 5.5, and is still pretty good. My take on the pricing, the cache-pricing gap, how it compares with GPT-6 Sol, and why it isn't my first pick for hard coding.

  • AI
  • Claude
  • LLMs
The words Far cheaper than Opus. Still pretty good. beside a bar chart of API prices per million tokens: Opus 5.5 at $4 input and $20 output, Sonnet 5.5 at $2 input and $10 output
Per token, Sonnet 5.5 costs half as much as Opus 5.5.

When Sonnet 5 launched in June, I wrote that it looked cheap at the token level but wasn't always cheap at the task level. I called it the junior engineer problem: a lower hourly rate doesn't help much if the work takes ten times longer.

Claude Sonnet 5.5, released on September 28, is Anthropic's answer to that problem. Mostly.

Most headlines skip this part: Sonnet 5.5 costs exactly the same per token as Sonnet 5. It's $2 per million input tokens and $10 per million output tokens. The tokenizer, cache prices and batch discount are all the same. If you were hoping for a smaller number on the pricing page, it isn't there.

What changed is how many tokens it burns to get the job done. That turns out to matter much more than the sticker price.

I've used it, gone through the benchmarks, and read the most useful outside review I could find. The best thing I found: next to Opus 5.5 it's very cheap, and it's still pretty good. That makes it an excellent general-purpose workhorse for the money. It isn't the model I'd reach for first when the coding gets hard, and Anthropic's pricing structure still has one gap I really wanted them to close.

Sonnet 5.5 at a glance

API price, input / output
$2 / $10
Per million tokens, unchanged from Sonnet 5
Intelligence Index, max effort
56
18 points above Sonnet 5, 2 below Opus 5.5
Cost per task, medium effort
$0.59
Under half of Opus 5.5's $1.34, and 41% less than Sonnet 5
Cache read price
$0.20
Per million tokens, the same as Opus 5.5

Released September 28, 2026 as claude-sonnet-5-5, with a 1M-token context window and up to 128K tokens of output. Index score and cost per task from Artificial Analysis; prices from Anthropic.

What Anthropic actually shipped

Sonnet 5.5 is the second model in the Claude 5.5 family, after Opus 5.5. Anthropic pitches it as the faster, cheaper complement to Opus. They say it's strongest at well-scoped everyday tasks, fixing bugs, and producing polished documents, slides and spreadsheets. It's available on Anthropic's API, AWS, Google Cloud and Microsoft's cloud.

Two things in the launch stood out to me.

First, Anthropic is unusually direct that Opus 5.5 is still clearly stronger at complex, open-ended work. Remember that line. It matters later.

Second, this is the first Sonnet to launch with Anthropic's cyber safeguards. Some higher-risk security requests will visibly fall back to Sonnet 5. If you're building security tooling, test your real workflows before you switch.

The pricing: what changed and what didn't

Per million tokensSonnet 5Sonnet 5.5
Input$2.00$2.00
Output$10.00$10.00
Cache write (5 min / 1 hr)$2.50 / $4.00$2.50 / $4.00
Cache read$0.20$0.20
Batch input / output$1.00 / $5.00$1.00 / $5.00
TokenizerSameSame

There's some backstory. Sonnet 5 launched at $2/$10 as introductory pricing through August 31, with an increase to $3/$15 scheduled for September 1. In August, Anthropic cancelled that increase and made $2/$10 the standard price. So if your budget still assumes $3/$15, the Sonnet line is a third cheaper than planned. But that happened to Sonnet 5 before 5.5 existed. Sonnet 5.5 didn't move the price again.

So where does "cheaper" come from?

Anthropic says Sonnet 5.5 costs up to 30% less per task because it uses fewer tokens and fewer tool calls. Vendors always say things like that, so I looked at independent numbers.

Artificial Analysis runs every model through the same set of evaluations and reports what one task costs at each effort level:

Same price per token, a very different bill

What one task on Artificial Analysis's Intelligence Index costs at each effort level. Sonnet 5.5 is cheaper at every level except max, and scores higher at all of them.

  • Sonnet 5
  • Sonnet 5.5

Low

Medium

High

Xhigh

Max

Source: Artificial Analysis, as of September 29, 2026. Hover or tap a bar for its Intelligence Index score.

Show the data
EffortSonnet 5 costSonnet 5 scoreSonnet 5.5 costSonnet 5.5 scoreSonnet 5.5 against Sonnet 5
Low$0.5124$0.413620% cheaper
Medium$1.0028$0.594141% cheaper
High$1.7932$1.084740% cheaper
Xhigh$2.8734$2.74525% cheaper
Max$5.0938$7.605649% more

This is the clearest example I can give. Sonnet 5.5 at medium effort scores 41 for $0.59 per task. Sonnet 5 at max effort scores 38 for $5.09. You get a better result for about 12% of the cost.

At high effort, 1,000 of these tasks cost about $1,790 on Sonnet 5 and $1,080 on Sonnet 5.5, while the score rises from 32 to 47. Same price list, a completely different bill.

Anthropic's launch customers report the same direction. Balyasny Asset Management said Sonnet 5.5 used about 121K tokens per answer where Sonnet 5 used 497K. Lovable reported a third fewer tool calls. CodeRabbit said Sonnet 5's "high token use" is gone. Anthropic picked these testimonials, so I weigh them less than independent data, but they agree with it.

The exception is max effort. Sonnet 5.5 at max costs 49% more per task than Sonnet 5 at max. It's far smarter (56 vs 38), but it isn't cheaper. Hold that thought.

My criticism: the cache pricing didn't move

In agentic work, a large share of what you pay for is cache reads: the model re-reading the same codebase, documents and conversation history on every turn.

With the 5.5 generation, Anthropic made cache reads much cheaper on its other models:

Anthropic cut cache-read prices on Opus and Fable, but not on Sonnet

Price per million tokens read from the prompt cache, previous model to current one.

  • Previous model
  • Opus 5.5, Fable 5.1
  • Sonnet 5.5

OpusOpus 5 → Opus 5.5

FableFable 5 → Fable 5.1

SonnetSonnet 5 → Sonnet 5.5

Source: Anthropic API pricing. Cache reads cost 5% of the input price on Opus 5.5 and 2.5% on Fable 5.1; Sonnet 5.5 stays at the standard 10%.

Show the data
ModelCache read per 1M tokensShare of input price
Opus 5$0.5010%
Opus 5.5$0.205%
Fable 5$1.0010%
Fable 5.1$0.252.5%
Sonnet 5$0.2010%
Sonnet 5.5$0.2010%

So reading a cached token now costs exactly the same on Sonnet 5.5 as on Opus 5.5, even though Opus charges twice as much for input and output.

Here's what that does to an illustrative agent task that reads 2M cached tokens across its loop, sends 100K fresh input tokens and writes 50K output tokens:

One agent task: where the money goes

An illustrative agent loop that reads 2M cached tokens, sends 100K fresh input tokens and writes 50K output tokens. Opus 5.5 costs 1.6 times as much here, not twice as much.

  • Cache reads (2M tokens)
  • Fresh input (100K tokens)
  • Output (50K tokens)

Sonnet 5.5Today's prices

$1.10

Opus 5.5Today's prices

$1.80

Sonnet 5.5If it had Opus 5.5's 5% cache rate

$0.90

An illustrative example priced at Anthropic's API rates, not a benchmark. Cache writes are left out to keep it simple.

Show the data
ScenarioCache readsFresh inputOutputTotal
Sonnet 5.5, today's prices$0.40$0.20$0.50$1.10
Opus 5.5, today's prices$0.40$0.40$1.00$1.80
Sonnet 5.5, if it had opus 5.5's 5% cache rate$0.20$0.20$0.50$0.90

On paper Sonnet is half the price of Opus. In this example, Opus costs only 1.6 times as much. If Sonnet needs more output tokens than Opus for the same job, which the benchmarks suggest can happen, the gap shrinks further.

That's the change I wanted to see. I didn't need a headline cut to input pricing. I wanted a cache structure that fits how people actually use a Sonnet-class model: long agent loops, sub-agents, and huge cached contexts. Sonnet is the model you'd run inside those loops, and it's the one that didn't get the cache discount.

The best part: much cheaper than Opus, and pretty good

This is what stood out most when I used it. Sonnet 5.5 costs a fraction of what Opus 5.5 does, and for most of what I've used it for, it's pretty good.

The numbers agree. Per token it's half Opus 5.5's price: $2/$10 against $4/$20. Per task, at the same effort level, Artificial Analysis measured it cheaper at every setting except max:

At the same effort level, Sonnet 5.5 costs far less than Opus 5.5

What one task on Artificial Analysis's Intelligence Index costs on each model at each effort level. Sonnet 5.5 is cheaper at every level except max; Opus 5.5 scores higher at all of them.

  • Opus 5.5
  • Sonnet 5.5

Low

Medium

High

Xhigh

Max

Source: Artificial Analysis, as of September 29, 2026. Hover or tap a bar for its Intelligence Index score.

Show the data
EffortOpus 5.5 costOpus 5.5 scoreSonnet 5.5 costSonnet 5.5 scoreSonnet 5.5 against Opus 5.5
Low$0.5542$0.413625% cheaper
Medium$1.3451$0.594156% cheaper
High$1.8254$1.084741% cheaper
Xhigh$3.4656$2.745221% cheaper
Max$5.9858$7.605627% more

At medium effort, a task costs $0.59 on Sonnet 5.5 and $1.34 on Opus 5.5, less than half. Even in the cache-heavy example above, where the gap is narrowest, Sonnet still costs about 40% less. Theo's review (more on that below) found a similar gap on a real codebase audit: about half of Opus's cost, with a slightly better score.

That's the appeal for me. For everyday work, "pretty good at less than half the price" beats "a bit better at twice the price" most of the time.

It's faster, too:

Output speed at high effort

Tokens generated per second. Speed is one place Sonnet 5.5 clearly beats Opus 5.5.

Sonnet 5

Sonnet 5.5

Opus 5.5

GPT-6 Sol

Source: Artificial Analysis, as of September 29, 2026.

Show the data
ModelOutput tokens per second
Sonnet 565
Sonnet 5.589
Opus 5.575
GPT-6 Sol71

The catch: Opus is still smarter, and sometimes the better deal

Opus 5.5 got cheaper as well, dropping from $5/$25 on Opus 5 to $4/$20, and it scores higher at every effort level. If you compare by result instead of by setting, the picture changes:

  • Opus 5.5 at low effort (42, $0.55) slightly beats Sonnet 5.5 at medium (41, $0.59).
  • Opus 5.5 at high (54, $1.82) beats Sonnet 5.5 at xhigh (52, $2.74).
  • Opus 5.5 at xhigh matches Sonnet 5.5 at max (56) for less than half the cost ($3.46 vs $7.60).

Sonnet 5.5 leaps past Sonnet 5, but Opus 5.5 sits higher for similar money

Intelligence Index score against the cost of one task. Each dot is an effort level, from low on the left to max on the right, so higher and further left is better value.

  • Sonnet 5
  • Sonnet 5.5
  • Opus 5.5

Intelligence Index

Cost per task

Source: Artificial Analysis, as of September 29, 2026. The cost axis is logarithmic.

Show the data
ModelEffortCost per taskIndex score
Sonnet 5Low$0.5124
Sonnet 5Medium$1.0028
Sonnet 5High$1.7932
Sonnet 5Xhigh$2.8734
Sonnet 5Max$5.0938
Sonnet 5.5Low$0.4136
Sonnet 5.5Medium$0.5941
Sonnet 5.5High$1.0847
Sonnet 5.5Xhigh$2.7452
Sonnet 5.5Max$7.6056
Opus 5.5Low$0.5542
Opus 5.5Medium$1.3451
Opus 5.5High$1.8254
Opus 5.5Xhigh$3.4656
Opus 5.5Max$5.9858

So "much cheaper than Opus" is true when you run both at the same setting, and that's why Sonnet 5.5 is my default. But if you need Opus-level results, Opus at a lower effort level gets you there for less than pushing Sonnet to its limit. And at max effort, Sonnet actually costs more than Opus.

One caveat on all of this: one benchmark suite has its own mix of tasks, and your workload won't match it exactly.

My take: a genuinely good general-purpose model

I like Sonnet 5.5. For everyday work it's the most useful Sonnet Anthropic has shipped. That covers writing, research, analysis, working through long documents, and answering questions about a codebase.

The numbers back that up. On the knowledge-work results in Anthropic's launch post, it's essentially level with Opus 5.5, and on computer use it's close behind (80.1% to 81.8% on OSWorld 2.1):

On knowledge work, Sonnet 5.5 is level with Opus 5.5

Elo-style ratings on two knowledge-work benchmarks, as Anthropic reported them at launch. Higher is better.

  • Sonnet 5
  • Sonnet 5.5
  • Opus 5.5
  • GPT-6 Sol

GDPval-AA v2.1

AA-Briefcase v1.1

Source: Anthropic's Sonnet 5.5 launch post, which footnotes its GPT-6 Sol figures.

Show the data
BenchmarkSonnet 5Sonnet 5.5Opus 5.5GPT-6 Sol
GDPval-AA v2.11449184418461487
AA-Briefcase v1.11359181118221483

Anthropic also says it writes more clearly than the previous generation. That's one of the few launch claims the outside review I read agreed with.

It isn't perfect. On Artificial Analysis's factual-knowledge test (AA-Omniscience), Sonnet 5.5 scored 54% to Opus 5.5's 66%. For knowledge-heavy questions I can't easily check, I'd still verify the answer or use Opus.

Coding: very good, but not my first pick

I don't think Sonnet 5.5 is the best model for coding. It's a good coding model, and a big step up from Sonnet 5. But when a task is hard, ambiguous or involves design, I'd still reach for Opus 5.5 or Fable 5.1. That's my opinion. Here's the evidence on both sides.

Coding benchmarks, as Anthropic reports them

Sonnet 5.5 leads on terminal work. On the harder coding tests it trails Opus 5.5, and on FrontierCode it trails GPT-6 Sol too.

  • Sonnet 5
  • Sonnet 5.5
  • Opus 5.5
  • GPT-6 Sol

Terminal-Bench 4.0

CursorBench 4.0

FrontierCode 1.1 (max)

Source: Anthropic's Sonnet 5.5 launch post. Anthropic didn't report GPT-6 Sol on Terminal-Bench or CursorBench, and footnotes some figures.

Show the data
BenchmarkSonnet 5Sonnet 5.5Opus 5.5GPT-6 Sol
Terminal-Bench 4.010.3%70.6%66.4%Not reported
CursorBench 4.034.1%55.5%57.8%Not reported
FrontierCode 1.1 (max)42.4%46.2%54.4%49.3%

Where it's strong

  • Terminal-style agentic work. On Terminal-Bench 4.0, Anthropic reports 70.6%, up from Sonnet 5's 10.3% and ahead of Opus 5.5's 66.4%. Artificial Analysis's own run agrees on the ranking (below).
  • Bug fixes and well-scoped tasks. This is exactly the area Anthropic says it's best at.
  • Efficiency compared with Sonnet 5. It uses fewer tool calls, fewer shell runs and fewer wasted steps.

Terminal-Bench 4.0, measured independently

Artificial Analysis's own run puts Sonnet 5.5 ahead of Opus 5.5 and GPT-6 Astra on command-line work, the same order Anthropic reported.

Sonnet 5

Sonnet 5.5

Opus 5.5

GPT-6 Sol

GPT-6 Astra

Source: Artificial Analysis, as of September 29, 2026.

Show the data
ModelTerminal-Bench 4.0
Sonnet 514.1%
Sonnet 5.563.6%
Opus 5.559.6%
GPT-6 Sol43%
GPT-6 Astra59.1%

Where it falls short

  • Harder coding benchmarks. It trails Opus 5.5 on CursorBench 4.0 and FrontierCode 1.1, and GPT-6 Sol on FrontierCode at max effort. Those are Anthropic's own published numbers.
  • More thinking isn't always better. In the same table, Sonnet scores higher on FrontierCode at xhigh (52.1%) than at max (46.2%).
  • Token appetite. Artificial Analysis measured about 193K output tokens per task at max effort, the highest it had seen. GPT-6 Sol used about 31K.
  • Frontend and design. Anthropic says it has a sharp eye for design. The outside review I read found its layouts clearly worse than Opus 5.5 and Fable 5.1.
  • Checking its own work at low effort. Anthropic's migration guide says that at low effort, Sonnet 5.5 sometimes reports a code change as done without running a check that exercises it. The guide includes a system-prompt paragraph to fix this. I appreciate the honesty, but it's the kind of thing you want to know before running it unattended.

General intelligence and coding performance are different things. GDPval-style benchmarks reward knowledge work. Hard coding rewards sustained judgment across many files and many decisions, and that's where Opus's advantage shows up.

The outside review worth watching: "delegate, don't select"

The most useful outside take I've found is Theo Browne's review on the T3 (t3.gg) channel, published the day after launch. It's useful partly because it doesn't simply accept Anthropic's framing.

Theo's central argument is that Sonnet 5.5 is a poor model to select as your main model and a great one to delegate to. In short, Sonnet is the model you hand work to, not the one you talk to.

The findings that stood out to me:

  • Theo's own benchmark: a deep pull-request audit of a large real-world codebase. Sonnet 5.5 cost about half as much as Opus, scored slightly better on the automated judging, and finished in about five minutes. Opus took roughly twice as long and GPT-6 Astra roughly three times as long.
  • Everyday coding cost. For like-for-like coding work, Theo found Sonnet often ends up costing as much as Opus, or more, because of cache reads and heavy token use. That isn't what I've seen day to day, but it's a fair warning for long agent loops, and it fits the cache example above.
  • Max effort. Token usage jumps dramatically from xhigh to max for very little gain. Theo's advice is to pretend max doesn't exist.
  • Cache pricing. Theo made the same criticism I did.
  • Design. Theo found it worse than Opus and Fable 5.1, though ahead of OpenAI's and xAI's models.

Where I land compared with Theo: I agree on cache pricing, on max effort, and on Sonnet's value as a sub-agent. I'm more positive about picking it directly as a general-purpose model. Theo's review focuses on coding workflows, which is exactly where Opus's lead is biggest. Theo is also much harsher on GPT-6 Sol than I am.

Sonnet 5.5 vs GPT-6 Sol: same price, opposite approaches

OpenAI released GPT-6 Sol on September 22 as its mid-tier "workhorse" model, below the flagship GPT-6 Astra. Its API price is $2 input, $0.20 cached input, $10 output, identical to Sonnet 5.5 at standard context length.

That's why I find this matchup so interesting. Price is off the table, so the question becomes how each model spends your money. OpenAI also took the opposite pricing approach to Anthropic: Sol's price is half of what GPT-5.6 Sol cost ($4/$20).

Claude Sonnet 5.5GPT-6 Sol
Input / cached / output (per MTok)$2 / $0.20 / $10$2 / $0.20 / $10
Long-context pricingFull 1M window at standard ratesSeparate higher rate: $4 / $0.40 / $15
AA Intelligence Index (max effort)5648
AA cost per task (max effort)$7.60~$1.05
Output tokens per task (max)~193K~31K
Terminal-Bench 4.0 (AA)63.6%43%
FrontierCode 1.1, max (Anthropic-reported)46.2%49.3%
GDPval-AA v2.1 (Anthropic-reported)18441487

Sol is frugal. Sonnet 5.5 spends more to reach a higher ceiling:

Same price per token, opposite approaches

Sonnet 5.5 and GPT-6 Sol on the Intelligence Index, one dot per effort level. At about $1 a task they're level; below that Sol gets more for the money, and above it Sonnet 5.5 keeps climbing where Sol stops.

  • Sonnet 5.5
  • GPT-6 Sol

Intelligence Index

Cost per task

Source: Artificial Analysis, as of September 29, 2026. The cost axis is logarithmic.

Show the data
ModelEffortCost per taskIndex score
Sonnet 5.5Low$0.4136
Sonnet 5.5Medium$0.5941
Sonnet 5.5High$1.0847
Sonnet 5.5Xhigh$2.7452
Sonnet 5.5Max$7.6056
GPT-6 SolLow$0.1334
GPT-6 SolMedium$0.2540
GPT-6 SolHigh$0.3843
GPT-6 SolXhigh$0.5244
GPT-6 SolMax$1.0548

Output tokens per task, at max effort

The same price per token, but Sonnet 5.5 writes about six times as many tokens to finish a task. That's most of the cost gap.

Sonnet 5.5

GPT-6 Sol

Source: Artificial Analysis's launch analyses of Sonnet 5.5 and GPT-6 Sol. Approximate figures.

Show the data
ModelOutput tokens per task
Sonnet 5.5~193K
GPT-6 Sol~31K

For the cheapest acceptable answer at scale, Sol is a serious option. For a model that can go further when I let it, and especially for knowledge work, where Sonnet's lead is large, I'd pick Sonnet 5.5.

How I'd actually use Sonnet 5.5

  • Default for everyday work instead of Opus, at low or medium effort. It's far cheaper and pretty good.
  • As a sub-agent for codebase exploration, research, audits and scoped fixes, with Opus 5.5 or Fable 5.1 orchestrating.
  • Skip max effort. Xhigh is the highest I'd go.
  • Hard coding and frontend work: Opus 5.5 or Fable 5.1.
  • Cache-heavy agent loops: measure Sonnet against Opus on your own workload. The gap is smaller than the list prices suggest.
  • Coding agents at low effort: add Anthropic's verification paragraph to the system prompt.

The verdict

Sonnet 5.5 is the most cost-effective Sonnet Anthropic has shipped. That isn't because the price went down. It didn't. It's because the model wastes much less. That mostly fixes the junior engineer problem I wrote about with Sonnet 5.

The best part, for me, is how it compares with Opus 5.5: very cheap, and still pretty good. For most everyday work, that's the trade I want.

Three things stop it from being an automatic choice. Anthropic left Sonnet's cache pricing alone while cutting it everywhere else. When you need Opus-level results, Opus 5.5 at a lower effort level can be the better deal. And for coding, Sonnet 5.5 is good, but not the best.

It's a strong release, just not one that's right for every job.


A note on sources: benchmark numbers labelled "Anthropic-reported" come from Anthropic's launch post. Index scores, cost per task, speed and token counts come from Artificial Analysis's independent testing. Prices come from Anthropic and OpenAI. Theo's findings are from the T3 review. Everything written as "I think" is my own opinion.