293 comments
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.
How do the numbers pan out? Ack that it isn’t always juts about cost, so even if it is pricier to self/host it might still be better for you for other reasons.
This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.
Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?
Sorry if it’s a dumb question, I don’t really know much about the topic.
GPT5.6-Sol on Max thinking just became regarded as of a few days ago.
The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).
The cycle repeats.
We are being A/B tested on and there is nothing you can do about it.
Thanks for your insight
Surely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.
Change is fast and abundant, and at the same time, it is hilariously more mundane than the dangerous-AGI-in-six-months tune we've been reading daily for years.
I would define it as a quick-moving market, but not nearly moving enough for the fantastic claims they make to justify ever-increasing funding.
I’m not saying it is stagnant. I’m saying for a hypothetical industry that was (maybe that fits AI, maybe not, I have zero authority to say myself)…
The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).
For my use cases, we are definitely on the flatter part of the curve at the moment.
AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.
There are several projects that repeat benchmarks on published models. None has ever found significant fluctations
Here's one example https://marginlab.ai/trackers/claude-code/
Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work.
This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.
If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.
A properly run OWM enables you to stop wondering if you're being ripped off and that's what AI needs because I think it's incredibly easy to just assume we're being ripped off because these companies are all built on a foundation of wonton theft. (Not that I really care about that — I think all information should be free, but still.)
> 12. General terms
> Changes to the Services. Our Services are novel and will change. We may sometimes add or remove features, increase or decrease capacity limits, offer new Services, or stop offering certain Services.
> Unless we specifically agree otherwise in a separate agreement with you, we reserve the right to modify, suspend, or discontinue the Services or your access to the Services, in whole or in part, at any time without notice to you. Although we will strive to provide you with reasonable advance notice if we stop offering a Service, there may be urgent situations—such as preventing abuse, responding to legal requirements, or addressing security and operability issues—where providing advance notice is not feasible. We will not be liable for any change to or any suspension or discontinuation of the Services or your access to them.
You're not buying a gallon of milk or a pound of flour. You're buying hosted software that the host reserves the right to modify.
The only regulation that we need right now is the model that's on tap
Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."
The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.
Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.
Self run infrastructure won’t have this cost but you have to manage the capacity and rollouts yourself, at which point it’s more obvious what’s happening, but the effects will be the same. The not knowing makes it harder, but also harder to plan your own work around.
That we have to guess at this is by far the worst part of the AI era. It feels like a dark cloud over my productivity. It makes my body tense for the entire day when it happens. Not healthy.
But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment.
Honestly I lost patience with Anthropic both clearly messing around with things like this and their agitation over regulation. They aren't good actors, and quite why so many blindly trust them with their company crown jewels is a mystery.
I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.
If you have a bank of rigorously tested benchmarks that you run every few days, with enough trials to know what your standard deviation is, and you are getting significant trends over time with those, that would be interesting.
But "I have a feeling" and "Seems like" really isn't a reliable signal at all, humans just can't handle perceiving these things reliably. On top of that changes in your work environment can easily pollute LLMs and change quality of results. Are things getting added to your memory or claude.md files that you don't realize? Is your project growing in size and thus claude is performing worse as more context is needed to work with it? etc etc
The implication is that humans are unreliable and shouldn't be trusted.
Or humans have certain shorthands when they complain on reddit, but their diagnoses are accurate for the specific context? If my AI does something stupid, am I not allowed to call it out? A NS-solving AI is still capable of not satisfying the abstract thing called the user experience. People have intelligent thoughts without compiling to lean.
OK, you say. Then let's get an aggregate benchmark for "intelligence". That doesn't prove that AI didn't flounder a specific use case that the user requested.
Classic moves: Humans are unreliable, converge to some "objective" benchmark that necessarily will quotient out the special cases, etc. Wonder how we'll be solving these issues in the AGI era - well, if you have an AGI that just replicates itself, dominates everybody because it's a machine and humans are soft fleshy creatures, and agrees with itself, fine. But part of the beauty of human experience is the messy part, and providing value is in the messy part.
The need to have a mental measure of competence for your fellow man is, most likely, a pre-human skill, probably with a dedicated bit of neurons for it. I think the problem is that those instincts were co-evolved with our fellow man, and, as you say, don't apply at all to a more non-deterministic system that, fundamentally, lacks some logic faculties that even small children have (simple riddle modifications, car wash question, etc).
At the end of the day the highest quality of benchmark tells you nothing if the man behind the curtain is constantly changing variables on you. You have no idea if you're really testing the same thing at all. So when you run your test at the top level on their system you're seeing lets say a 30% difference in quality most of the time, you have no idea if you should really only see a 5% difference in quality if you were running a local model with stable settings.
Opus 5.5 is being served under opus 5 right now.
On what basis are you claiming this?
For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"
I had to tell it to ssh into the server and run journlctl to check it
Anecdotal, I know, but they all seem to be less capable with time.
_edit_ I use the same reasoning level of `medium`
As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par.
Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money.
Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.
If for example AWS sells you a 2xLarge server there may be some variability in performance but it's going to be averaged out very well.
When it comes to AI services executing your model there is absolutely no information on what and with what settings your model is being executed. Hell, you have no idea if it even is the model you're paying for. Add that models are not deterministic so variability can be pretty large.
This leads to a common set of dynamics that induce cheating behavior in humans. For example, is there a mix of different hardware. Does lessor hardware use different settings? How do you know xhigh is what your prompt ran under. Anthropic has a proven history of running your prompt silently under different models.
This is a huge mess that needs and will be regulated or sued heavily over. Hell, with as many people out there that hate AI it might be easier than one thinks to have a state sue the providers on this and elicit a huge amount of discovery.
I wonder what their official explanation for this behavior is.
(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...
But it's still happening: https://github.com/anthropics/claude-code/issues/81759
This still leaves an absurd amount of wiggle room for arguments like "oh no, our evals show that this quantization has no detectable effect on performance (in the eval distribution) therefore running the quant doesn't degrade quality"
At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.
Read the full thread on Hacker News →
Related stories
- Hacker News · 10 points · 4 days ago
- First Principles Thinkingsunilsadasivan.comHacker News · 282 points · 5 days ago
- Hacker News · 2 points · 8 days ago
- First Principles Thinkingsunilsadasivan.comHacker News · 2 points · 7 days ago
- DEV Community · 6 points · 9 days ago
- Hacker News · 2 points · 6 days ago