Need help?
<- Back

Comments (141)

  • jug
    We also have Nerf Bench:https://www.bridgebench.ai/nerf-benchThey test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
  • sheepscreek
    > It could also mean nothing happened and people are pattern-matching on noise.It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
  • johnfn
    "Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.I made a graphic to explain why people feel like the models get nerfed:https://x.com/thesilenceturns/status/2103551351825543610The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
  • nico
    Anecdata: I've been running a long-lived claude code session with Opus 4.6 for the last few days. Yesterday, almost right after the Sonnet 5.5 announcement, codex starting asking for permission to run things a lot more oftenThe quality of the output/work seems the same, but the speed at which it gets stuff done is a lot slower, because it's asking for permission so much moreI don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output
  • xlayn
    The only reason why claude fable is better than opus in my opinion is that it has more "criteria"... if you present a problem and then ask for his recommendation you can get an opinion on why and reasoning on why that one... Opus is going to vomit 10k lines of extremely dense prose in nerdify++ level.Yesterday I fought claude fable to not just jump to make changes like a dog following a treat, that we were researching... at some point I introduced the word HAWAI... and only if I say HAWAI the thing can start making changes..I was going to post here in HN just to have a "I knew this was the reason" when they release fable > 5.1I had the exact same feeling every time they have a new big release
  • judge2020
    I wonder if more organizations approving the model on a fast-tracked basis means Anthropic is straining for more compute and thus sheds a tiny bit to handle the increased demand, especially at peak times.
  • gr_norm
    All this dishonesty and shadiness is part of why open models feel inevitable. Even if the total cost of ownership is higher (debatable; seems that way at small scales, but likely not as you grow), I'd rather have intelligence controlled by me that works for me.The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.
  • zeroonetwothree
    If it can’t even tell apart Opus 5 and 5.5 (according to the readme) then it’s not useful
  • solfox
    It seems as if this is based on demand. Whenever a new model is released, I'm guessing tens of thousands of us switch over to try the latest and greatest, which overloads the servers, leading to nerfing. It's 100% dishonest, but they realized they would lose users a lot quicker if they were honest and just said "our models are overloaded, come back later".After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?
  • winwang
    Complete anecdote, and nothing to do with relative nerfing or not: Opus 5.5 has been surprisingly good for me (including the past couple hours), especially for following research-level questions/directions.
  • aabhay
    Only ten day interval? I felt Astra got nerfed within a week
  • whs
    I wonder if API is affected by this issue, especially Claude on public clouds? Would that means the subsidized rate just means they use cheaper quantized models and it's not comparable to API spending.
  • gaigalas
    The nerfing/quantization strategy is unsustainable. The first lab to not do it wins (short term). The Anthropic pause on Fable might just have been that.My gut tells me this involves an undisclosed, never-released grandparent model (higher-class than Fable/Astra level, roughly unsellable due to unfeasible cost). That grandparent model is distilled into lower models, of which Opus 5.5 might be an instance of.That also guarantees protection against distilling a core business. You never make your prime weights available to the public, you only make distillings themselves available.The downside of this strategy is that you spend a lot of compute on something that you never release, but it might be just the right play (for now) for closed weight companies.It's a gut feeling, I have zero hard evidence to back it up (it's what I would do as them).
  • LeoPanthera
    n=1 is useless. The output is not deterministic.
  • solenoid0937
    Hot take, none of the models are getting "nerfed", people are just getting used to the new level of intelligence.
  • anon
    undefined
  • apt-apt-apt-apt
    Fable 5 seems like it got nerfed when 5.1 came out.
  • avazhi
    This explains a lot actually.First two days of this thing was like working with Einstein, then about 24-36 hours ago I started getting frustrated at bullshit that hadn't been a problem before. It was so egregious that I checked to make sure I was still on Opus 5.5 Max.
  • octoberfranklin
    OpenAI will simply set up a classifier to detect if the client is livenerf, and selectively not nerf those requests.Open models are the endgame.
  • lqstuart
    You used Claude to make some slop to see if Claude is getting worse…?
  • onlyrealcuzzo
    This is bad data at its finest.Truly, madly, deeply sloppy.
  • j45
    New model releases that have positive reviews should come with a nerfalert reminder service to make hay until it's shaped and shaped and shaped.
  • anon
    undefined
  • colordrops
    This repo already has too much visibility now. Anthropic will soon benchmaxx it.
  • gigatexal
    This is genius. I’m so worried opus 5.5 will get nerfed cuz sonnet 5 was such trash I can’t go back.
  • bethekidyouwant
    People just tend towards conspiracies you have to actively fight it.
  • folayii
    [dead]