Need help?
<- Back

Comments (196)

  • simonw
    Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.https://tools.simonwillison.net/markdown-svg-renderer?url=ht...Here's how the thinking effort levels compare: low 27 input, 1,623 output, thinking_tokens: 0 1.6284 Duration: 10138ms (10s) medium 27 input, 1,796 output, thinking_tokens: 0 1.7914 cents Duration: 11266ms (11s) high 27 input, 2,334 output, thinking_tokens: 745 2.3394 cents Duration: 17376ms (17s) xhigh 27 input, 5,730 output, thinking_tokens: 2535 5.7354 cents Duration: 41882ms (41s) max (failed to return response) 27 input, 128,000 output, thinking_tokens: 128000 $1.28 Duration: 940617ms (15m 40s) Low and medium both used 0 thinking tokens.
  • Sol-
    Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
  • MisterMunchkin
    It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.And my job won’t even pay for Claude now because it’s so ruinously expensive.
  • abejora
    Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.[1] Section 8.5 of the Sonnet 5.5 System Card
  • __jl__
    artificialanalysis.ai benchmarks are [here](https://artificialanalysis.ai/articles/claude-sonnet-5-5). Anthropic is back at spot 1, 2, 3 and 5. Impressive even if these benchmarks are problematic in many ways.
  • wongarsu
    "Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
  • ghoshbishakh
    So sonnet is better than Fable now? That Fable which was too dangerous to release? I am so confused now.
  • johnmlussier
    Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.This is bollocks. Their safeguards are shit.
  • wkcheng
    The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
  • heyjstn
    Have anyone tried a workflow that:- Fable 5.1 for planning/adversarial reviewer- Opus 5.5 for well-scoped tasks break down- Sonnet 5.5 for these well-scoped tasks implementationI think the blocker might be how efficient the context is compacted and sending around between these agents
  • Jcampuzano2
    I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
  • a13o
    This doesn’t have an interesting footprint on the intelligence/cost Pareto line compared to existing Opus 5.5 and GPT-6 models.
  • nanook
    Sonnet is 1/5th the price and seemingly more powerful than fable (the model that was too powerful to release). I can't make sense of this. Why would anyone use fable now? Or are the benchmarks completely pointless and one has to just try em to get a feel for what they can and can't do?
  • sajithdilshan
    I use Claude Code everyday for work and the main model I use is Opus (For planning, breaking down tasks, writing tickets, implementation, etc.) and Haiku for running tests. Honestly have no idea what is the use case for Sonnet
  • avree
    Crazy bad front-end design. Site hijacks my gestures so I can't swipe back anymore, starts with a full page autoplaying video...
  • yapfrog
    From the graph it looks like I'd rather use Opus 5.5 High than Sonnet 5.5 at all
  • alansaber
    Always key to include the one bench where the smaller model inexplicably outperforms the larger model
  • gregwebs
    This is better priced than Opus for tasks that are token heavy but not complicated. But a quick look shows that at least on some benchmarks DeepSeek performs as well and of course the cost is an order of magnitude less.From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.
  • tombert
    I like that "alignment on safety" appears to mean, at least for anything I've been doing, that they won't violate Microsoft's terms of service. I even had it pushing back on me activating an LTSC key on Windows because LTSC keys are "often purchased on a gray market and violate Microsoft's TOS".
  • dom96
    I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions whether to keep going.https://bench.killswitch-lang.org/ Claude Sonnet 5 17.8% Claude Sonnet 5.5 7.4%
  • onlyrealcuzzo
    > In our testing, it costs up to 30% less per task than its predecessor.> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.Hopefully they release a Haiku that actually has a reason for existing.
  • mroche
    Is there ever any focus on producing new Haiku models? There are a lot of use cases for quick to return models when you're limited to a single provider.
  • AM1010101
    For me I would like to pair this with Opus 5.5 as orchestrater and use Sonnet as a sub agent. Therefore I want it to be fast when on low or medium and not break the bank.On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.
  • ChickeNES
    Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".
  • s3p
    I'm loving the tit for tat cost charts these guys are doing. Just a few days ago it looked like OpenAI ruled the cost pareto frontier. Not even a week later and Anthropic is taking the charts again. See you guys same time next week?
  • s314
    In the Artificial Analysis Intelligence Index, Claude Sonnet 5.5 is the second best model behind Opus 5.5. This however is with max effort which costs even more than Opus 5.5 max. But Sonnet 5.5 xhigh is cheaper than Opus 5.5 xigh and matches GPT 6 Astra xhigh in the benchmark.
  • pookieinc
    It's interesting that in all their benchmarks, they omit Fable numbers and only focus on Opus, Sonnet, and OpenAI models. Maybe Fable is out the door?
  • swingboy
    Is Opus still 2x usage of Sonnet after this? My Claude Code isn't showing that warning anymore when I look at /model.
  • alasano
    I wonder if Fable 5.5 is coming this week to drown out the OpenAI dev day announcements
  • solenoid0937
    Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
  • system2
    Make 1M tokens $0.10; then I will use Sonnet. Until then, it is garbage.
  • taurath
    After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
  • croemer
    Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome.
  • ghoshbishakh
    So Sonnet 5.5 on max effort is as expensive as Fable 5.1? Because it uses a ton of tokens for a task.In xhigh effort it is a lot cheaper and possibly lot less impressive?
  • pavitheran
    Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.
  • takerofnaps
    Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
  • bayesianbot
    Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird
  • square_usual
    Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
  • rtuin
    Any benchmarks other than computer use/agentic coding published yet? Curious to compare more broadly with other models
  • limsungkee
    Yesterday, I realized that Opus 5.5 is cheaper than Sonnet 5. Now I know the reason.
  • _fw
    I still can’t find a place for Sonnet models, I never have.I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.Give me the frontier, or give me the cheapest form of good enough.
  • Alifatisk
    In other news> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
  • ramish94
    In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
  • laurenz-bauer
    Oh yes. I think you might get a lot for what you pay with Sonnet 5.5.
  • SeriousM
    Next will be haiku 5.5, surpassing opus 4.8
  • iagocc
    Waiting for the pelicans
  • jtrn
    Here's my purely academic initial impression based on only what they have released from the blog and the system card:If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.BUTIt regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.Some of the more interesting things I found from scanning the system card:- It is the only model tested that shows no preference for rude or polite style.- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).Clinical behaviour:Suicide and self-harm handling is reported as weaker in the API because itIt sometimes called a wish to die understandable.It sometimes validated self-harm as functional.It sometimes suggested harmful substitute behaviours.As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be
  • anon
    undefined
  • dude250711
    It's strange that there are no Astra comparisons. I guess they are positioning it as a Fable competitor. For me it's just a coding workhorse though, without any "fall-backs".
  • enraged_camel
    Another amazing release. This, combined with Opus 5.5, puts OpenAI in an incredibly tough spot: it means Anthropic's both mid-tier models crush OpenAI's top-tier model in capability and are also faster and significantly cheaper.If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...
  • dack
    very annoyed they aren't showing fable on the graph.
  • ahriad
    Time to switch team to Claude from OpenAI again.