Benchmark for tracking model capability after release. - ninjahawk/livenerf
377 comments
https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.
e.g. The original apple shuffle and the Risk app ins which a string of songs from the same album or three one roles are "not random"
`percieved_performance = actual_perf/expectation`
`expectation` is an increasing function over time.
`actual_perf` is a stochastic function of the model's true ability, context, etc. -> a recipe for some bad sessions.
As for multiple bad sessions in a row, this is a studied phenomenon in gambling where players perceive "runs" because our brains love to find patterns.
I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.
It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".
People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.
Anthropic cut a deal with SpaceXAI in May - $1.25b/mo. Before that, they employed months of dishonest nerfy strategies, to an extreme.
It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.
But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
What "stack" do you have in mind here?
An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.
Both Anthropic and OpenAI used to have multiple snapshots per model; both seem to have switched to updating version numbers when anything changes, looking at the "snapshots" section in their recent and old models: e.g. https://developers.openai.com/api/docs/models/gpt-4o for how far back I had to go to find multiple snapshots on an OpenAI model, and https://platform.claude.com/docs/en/about-claude/model-depre... seems to be a more useful list for Anthropic, but both are now things they did a year ago.
It most definitely is not, hasn't been for a while now.
I.e. when dealing with hosted models of the large providers, you are not interacting with a big bag of floats. You are interacting with an API/UI that presents an unholy web of software components, some of which may be large or small bags of floats, as if they were a big bag of floats.
Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
There's a lot of things to tune there, and just as many reasons to do it.
If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.
Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.
Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.
And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.
Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.
I don't think you understand what "this" is--or rather, the "thing" that didn't happen in "nothing happened".
> Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
So, not the sort of thing referred to.
> The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
Yes, but you claimed that this definitely did happen. But the whole point is to determine whether it did.
I don't follow. That sounds to me exactly like a reason why it could be people pattern-matching on noise: because there is a lot of noise in which a matchable pattern could emerge.
A few bad answers can be annoying as hell, but who knows what's behind them. I'm curious to see how the numbers look after a few weeks. Props for tracking improvements too.
Now I dismiss it every time and the quality is more consistent.
Complete adhoc and personal experience but something I've observed, wouldn't be surprised if they nerfed on a per session basis
Random performance is random, your brain will jump through hurdles to fit patterns where there aren't any.
In fact, since they have some rule based system (if bioenegineering or security, route to degraded model) it would be almost trivial to add 'user has filed feedback' to it.
Not saying this is what happens, but just that it's not as insane as it sounds.
And strangely, expressing frustration multiple times in a row would reliably trigger a feedback popup as well.
That said, given the propensity for mature code bases to have "fuck" in commit messages/comments and those are typically of higher quality, I curse up a storm when the clankers make mistakes, if only to try put more quality-code valence into context. https://news.ycombinator.com/item?id=36584464
I made a graphic to explain why people feel like the models get nerfed:
https://x.com/thesilenceturns/status/2103551351825543610
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Your chart is wrong.
Where?
I am more skeptical about the compute provider claim - do you have any evidence of that?
Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.
The majority of people are not on it, and the links are gated by a ton of toxic dark patterns and horrible UX trying to force people to sign up or log in.
I try to click on the image to enlarge and make the text readable, and I'm greeted with a login screen instead of a larger image.
https://marginlab.ai/trackers/claude-code/
This site has been documenting it for a while
You can't just dismiss something backed by careful measurements by throwing a truism at it. What is this honeymoon phase? Can you quantify it? If not, how are you sure it's real?
Read the full thread on Hacker News →
Related stories
- Hacker News · 3 points · 8 days ago
- DEV Community · 0 points · about 9 hours ago
- Hacker News · 1 points · 3 days ago
- Hacker News · 1 points · 9 days ago
- DEV Community · 7 points · 3 days ago
- Hacker News · 1 points · 9 days ago