Benchmark for tracking model capability after release. - ninjahawk/livenerf

883 points•bryan0•1 day ago•377 comments•

377 comments

jug1 day ago
We also have Nerf Bench:

https://www.bridgebench.ai/nerf-bench

They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.

This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.

nsarrazinabout 18 hours ago
Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time.

It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.

goodmythicalabout 12 hours ago
This reminds me of the fact that true random does not feel random to users due to the clumpiness that the average person does not anticipate existing in true random.

e.g. The original apple shuffle and the Risk app ins which a string of songs from the same album or three one roles are "not random"

waterproofabout 17 hours ago
I find that I learn to "trust" a model to get certain things right, as I would trust a colleague. So, as `expectation` increases, my prompting and context management gets sloppier.

`percieved_performance = actual_perf/expectation`

`expectation` is an increasing function over time.

`actual_perf` is a stochastic function of the model's true ability, context, etc. -> a recipe for some bad sessions.

As for multiple bad sessions in a row, this is a studied phenomenon in gambling where players perceive "runs" because our brains love to find patterns.

causalabout 12 hours ago
Yeah I bet most of us remember GPT-4 a lot more fondly than we would if we were to return to it today.
rplntabout 20 hours ago
> I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.

I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.

It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".

transcriptaseabout 19 hours ago
I vividly remember when ChatGPT3.5 went fully mainstream, there were times where within minutes you would realize they were only serving up idiot mode and there was no point trying to do much until demand died down and they swapped back to the non-quantized version.

People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.

Barbingabout 11 hours ago
> It's been a few months since I last recall this though

Anthropic cut a deal with SpaceXAI in May - $1.25b/mo. Before that, they employed months of dishonest nerfy strategies, to an extreme.

https://www.anthropic.com/news/higher-limits-spacex

silversmithabout 16 hours ago
Anecdote - I work in a time zone offset from continental US. The performance of anthropic models would noticeably drop, around the time US work day started. It was so bad around 4.x time that multiple colleagues re-arranged their schedule to have least overlap with US work day. Admittedly it's been better recently.
smurf9852about 17 hours ago
Friend of mine works for a corp that is one of the top spenders on Claude models. He complained about these nerfs during peak demand. Their Anthropic contact changed something and it did not happen since.
comboy1 day ago
But you are using API not the CLI right? I did not ever observe API degradation, only subscription stuff through their CLI.
Aperockyabout 12 hours ago
So that's why even Qwen-3.8-27B caught up with Opus4.6, it was nerfed to the ground.
scrollopabout 20 hours ago
There's also this one which has been around for a while

https://marginlab.ai/trackers/claude-code/

sscaryterryabout 16 hours ago
The tracker for Codex resonates with me. Its thick as pig shit the last few days: https://marginlab.ai/trackers/codex/
sheepscreekabout 23 hours ago
> It could also mean nothing happened and people are pattern-matching on noise.

It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.

What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.

But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.

ben_wabout 19 hours ago
> It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

What "stack" do you have in mind here?

An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.

Both Anthropic and OpenAI used to have multiple snapshots per model; both seem to have switched to updating version numbers when anything changes, looking at the "snapshots" section in their recent and old models: e.g. https://developers.openai.com/api/docs/models/gpt-4o for how far back I had to go to find multiple snapshots on an OpenAI model, and https://platform.claude.com/docs/en/about-claude/model-depre... seems to be a more useful list for Anthropic, but both are now things they did a year ago.

TeMPOraLabout 19 hours ago
> An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.

It most definitely is not, hasn't been for a while now.

I.e. when dealing with hosted models of the large providers, you are not interacting with a big bag of floats. You are interacting with an API/UI that presents an unholy web of software components, some of which may be large or small bags of floats, as if they were a big bag of floats.

Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.

There's a lot of things to tune there, and just as many reasons to do it.

dannywabout 18 hours ago
Serving LLMs at Anthropic scale is very, very different. It’s not SGLang or vLLM.

If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.

Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.

Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.

And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.

Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.

jibaltabout 18 hours ago
> It is definitely not this.

I don't think you understand what "this" is--or rather, the "thing" that didn't happen in "nothing happened".

> Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

So, not the sort of thing referred to.

> The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.

Yes, but you claimed that this definitely did happen. But the whole point is to determine whether it did.

zahlmanabout 20 hours ago
> It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

I don't follow. That sounds to me exactly like a reason why it could be people pattern-matching on noise: because there is a lot of noise in which a matchable pattern could emerge.

ChristmasTomerabout 10 hours ago
Glad someone's keeping track. These threads always turn into “they made it dumber” vs “nah, you're imagining it.”

A few bad answers can be annoying as hell, but who knows what's behind them. I'm curious to see how the numbers look after a few weeks. Props for tracking improvements too.

msejasabout 19 hours ago
As a claude code power user, when I get the 'rate the feedback on Claude' pop up, I used to say good or fine out of habit, and immediately after sending this feedback, I felt an instant degradation and mistakes that usually don't happen.

Now I dismiss it every time and the quality is more consistent.

Complete adhoc and personal experience but something I've observed, wouldn't be surprised if they nerfed on a per session basis

xnorswapabout 15 hours ago
This is "rub your gameboy the right way to catch more pokemon" levels of insanity.

Random performance is random, your brain will jump through hurdles to fit patterns where there aren't any.

unglaublichabout 14 hours ago
On the one hand, yes, on the other hand it would also not be too hard to do routing based on such a parameter and Antrophic has shown (through Fable and Mythos) that they have such a transparent routing system in place.

In fact, since they have some rule based system (if bioenegineering or security, route to degraded model) it would be almost trivial to add 'user has filed feedback' to it.

Not saying this is what happens, but just that it's not as insane as it sounds.

Aurornisabout 13 hours ago
As a counterexample, I click that feedback all the time and nothing ever changes afterward.
SilverSlashabout 13 hours ago
My personal issue with clicking the feedback is that I'm doing free RLHF for Anthropic without getting anything in return.
itopaloglu83about 19 hours ago
Another personal anecdote, but I had the case where Claude would visibly improve after a negative feedback.

And strangely, expressing frustration multiple times in a row would reliably trigger a feedback popup as well.

okwhateverdudeabout 18 hours ago
When the source maps leaked for the harness, it was revealed that they track how often tool use is rejected and how many times you say "fuck". They definitely try to track frustration sentiment.

That said, given the propensity for mature code bases to have "fuck" in commit messages/comments and those are typically of higher quality, I curse up a storm when the clankers make mistakes, if only to try put more quality-code valence into context. https://news.ycombinator.com/item?id=36584464

Galilyouabout 19 hours ago
Wow .. really! Definitely testing that out a few times meself
johnfn1 day ago
"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

I made a graphic to explain why people feel like the models get nerfed:

https://x.com/thesilenceturns/status/2103551351825543610

The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

prodigycorp1 day ago
Incorrect.

Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.

Your chart is wrong.

simonw1 day ago
> Anthropic has admitted to nerfing in the past

Where?

johnfn1 day ago
Sorry, you are correct - I modified my original post. I get frustrated every time there's a model release and 1 week later everyone is saying NERF! NERF! 99.9% of the time these people are wrong, but you are right that it's technically not 100% due to a few edge cases.

I am more skeptical about the compute provider claim - do you have any evidence of that?

swader9991 day ago
Right, and it would be simple to un-nerf or shadow nerf by any kind of angle they want.
gobdovan1 day ago
There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs, e.g.: https://www.anthropic.com/engineering/april-23-postmortem

Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.

johnfn1 day ago
You are correct, I softened the wording a bit. And thanks for the heads up on the typo!
sspiffabout 16 hours ago
Off topic for sure, but why do people insist on using X/Twitter in this day and age?

The majority of people are not on it, and the links are gated by a ton of toxic dark patterns and horrible UX trying to force people to sign up or log in.

I try to click on the image to enlarge and make the text readable, and I'm greeted with a login screen instead of a larger image.

s08148692about 14 hours ago
There is an absolutely massive tech community on Twitter and its by far the place to get real time updates on tech news (yes - better than HN). It isn't all a far right cess pit and that is easy to avoid by just using the following tab
stratos123about 15 hours ago
Get an extension like LibRedirect and load it up with some public nitter instances and you can get an actually sane twitter-browsing experience.
scrollopabout 20 hours ago
It's a real thing

https://marginlab.ai/trackers/claude-code/

This site has been documenting it for a while

cbg0about 20 hours ago
You've posted a link that doesn't support your statement.
winwang1 day ago
That's a good observation, though I'd say here that two things could be true at the same time. But, I do personally believe that most of the reported nerfing is the case of your chart + latent evidence-less complaining. Honeymoon phases are real.
khalicabout 15 hours ago
> Honeymoon phases are real

You can't just dismiss something backed by careful measurements by throwing a truism at it. What is this honeymoon phase? Can you quantify it? If not, how are you sure it's real?

Read the full thread on Hacker News →

Related stories