425 points•espeed•9 days ago•293 comments•

293 comments

talon86359 days ago
Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.

AmazingTurtle9 days ago
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.

Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.

boardwaalk9 days ago
I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.
zeroonetwothree9 days ago
Last time I estimated it was like 30 years to pay back. I doubt the hardware will even last that long.
AtHeartEngineer9 days ago
flash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans
ramesh319 days ago
>Also I will likely save some money on subscriptions.

Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.

andsoitis9 days ago
I assume you’ve calculated expected cost vs subscription.

How do the numbers pan out? Ack that it isn’t always juts about cost, so even if it is pricier to self/host it might still be better for you for other reasons.

Aurornis9 days ago
> to create a perceived improvement when in reality there isn’t really one?

This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.

ruszki9 days ago
Overfitting to benchmarks. And puff, you have the exact same effect.
pixl979 days ago
Far more likely it's about reducing costs.
talon86359 days ago
This is a good point I hadn’t considered, thank you.

Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?

Sorry if it’s a dumb question, I don’t really know much about the topic.

trenchgun9 days ago
The public frontier is not the frontier. Actual frontier models are too expensive to serve to public, and also risk distillation by competitors.
AnimalMuppet9 days ago
Also, in a world where there are several models competing with each other for public perception of which is best, that seems like an extremely bad move.
nxc189 days ago
There must be some benefit if all the providers are doing it independently.

GPT5.6-Sol on Max thinking just became regarded as of a few days ago.

The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).

The cycle repeats.

claydugo9 days ago
Astra is also useless and completely ignoring instructions at random intervals.

We are being A/B tested on and there is nothing you can do about it.

talon86359 days ago
Again, I’m out of my element here, but isn’t the entire industry dependent on “new better releases frequently”? If so, and if no one has made any meaningful breakthrough, might they all pursue this kind of deception just to stay afloat/“competitive”/relevant?

Thanks for your insight

bitexploder9 days ago
They obviously test various quants and other serving cost saving strategies. Models like Fable are probably trillions parameters with hundreds of billions active MoE. They probably try to squeeze and quant each piece until people notice.
progval9 days ago
This sounds similar to rumors about how SSD companies work. First they would design a new drive with better performance that everyone uses to benchmark against other models; then slowly change its parts to worse ones, either because they are cheaper, the originals are no longer available, or whatever reason
bmicraft9 days ago
That's not a rumour, there are countless recorded examples.
a2dam9 days ago
> For an industry that’s stagnant in progress

Surely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.

Lalabadie9 days ago
I get the perspective from which you're making that statement, but the industry keeps moving its own goalposts.

Change is fast and abundant, and at the same time, it is hilariously more mundane than the dangerous-AGI-in-six-months tune we've been reading daily for years.

I would define it as a quick-moving market, but not nearly moving enough for the fantastic claims they make to justify ever-increasing funding.

talon86359 days ago
It’s a hypothetical statement that seems to have confused a lot of people.

I’m not saying it is stagnant. I’m saying for a hypothetical industry that was (maybe that fits AI, maybe not, I have zero authority to say myself)…

koyote9 days ago
And yet they have only improved marginally in my use cases since around Opus 4.5.

The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).

For my use cases, we are definitely on the flatter part of the curve at the moment.

jesse_dot_id9 days ago
The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.

AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.

Aurornis9 days ago
I doubt that would change the perception. Every model release is followed by accusations of nerfing.

There are several projects that repeat benchmarks on published models. None has ever found significant fluctations

Here's one example https://marginlab.ai/trackers/claude-code/

Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work.

This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.

ricardobeat9 days ago
This page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'.

If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.

jesse_dot_id9 days ago
It would change my perception but only if there were a competent and stringent administration in place. I didn't used to have to wonder if the ground beef I was buying was actually 1lb because there were inspections and repercussions, but stuff is kind of chronically underweight these days.

A properly run OWM enables you to stop wondering if you're being ripped off and that's what AI needs because I think it's incredibly easy to just assume we're being ripped off because these companies are all built on a foundation of wonton theft. (Not that I really care about that — I think all information should be free, but still.)

wongarsu9 days ago
If you go to https://marginlab.ai/trackers/claude-code-historical-perform... there is a very clear downwards trend in the two weeks before Opus 4.7 release. Then a sudden and dramatic drop seven days before Opus 4.8. And now we seem to have entered another decline in the last ten days, beyond the usual noise of Opus 5 scores
bradleybuda9 days ago
Anthropic terms of service:

> 12. General terms

> Changes to the Services. Our Services are novel and will change. We may sometimes add or remove features, increase or decrease capacity limits, offer new Services, or stop offering certain Services.

> Unless we specifically agree otherwise in a separate agreement with you, we reserve the right to modify, suspend, or discontinue the Services or your access to the Services, in whole or in part, at any time without notice to you. Although we will strive to provide you with reasonable advance notice if we stop offering a Service, there may be urgent situations—such as preventing abuse, responding to legal requirements, or addressing security and operability issues—where providing advance notice is not feasible. We will not be liable for any change to or any suspension or discontinuation of the Services or your access to them.

You're not buying a gallon of milk or a pound of flour. You're buying hosted software that the host reserves the right to modify.

winrid9 days ago
Companies can say whatever they want it doesn't mean we'll agree it's okay or legal.
alightsoul9 days ago
You are not buying something and expecting it to be what's on the tin? Aka what the benchmarks show?
moffkalast9 days ago
Petition to rename them to the Office of Weights and Biases, haha.
vatsachak9 days ago
THIS EXACTLY.

The only regulation that we need right now is the model that's on tap

lz4009 days ago
we could even just repurpose the same office, "weights and measures" is oddly relevant
Waterluvian9 days ago
I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.

Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."

The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.

Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.

fnordpiglet9 days ago
It’s load shedding. They’re reducing consumption for capacity balancing at your expense. Whenever there are rate limiting storms Claude gets dumber. They also shift capacity for new releases, and Claude gets dumber leading up to it.

Self run infrastructure won’t have this cost but you have to manage the capacity and rollouts yourself, at which point it’s more obvious what’s happening, but the effects will be the same. The not knowing makes it harder, but also harder to plan your own work around.

rfgplk9 days ago
Correct, but they should explicitly announce this ahead of time.
zarmin9 days ago
I would rather wait in a queue than be routed to a degraded model. And if they _have_ to degrade the models, then I wish they would fucking tell us. Instead, it's "I have a strong feeling".

That we have to guess at this is by far the worst part of the AI era. It feels like a dark cloud over my productivity. It makes my body tense for the entire day when it happens. Not healthy.

meowface9 days ago
They have repeatedly said they do not ever intentionally reduce model quality and do not degrade in this way, and that a model version number is always the same.

But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment.

fidotron9 days ago
The Claude models definitely felt more susceptible to moods, like you could leave them for a few hours, come back and it suddenly was unable to do things which it was doing just earlier, which tellingly is never an experience I've had with an open model.

Honestly I lost patience with Anthropic both clearly messing around with things like this and their agitation over regulation. They aren't good actors, and quite why so many blindly trust them with their company crown jewels is a mystery.

espeed9 days ago
Claude Code's prompt cache expires after 1 hour.
JMKH429 days ago
If you follow reddit forums for claude code, its common to see people, on the same day, claiming that Opus/Fable is especially smart today, and especially dumb today.

I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.

If you have a bank of rigorously tested benchmarks that you run every few days, with enough trials to know what your standard deviation is, and you are getting significant trends over time with those, that would be interesting.

But "I have a feeling" and "Seems like" really isn't a reliable signal at all, humans just can't handle perceiving these things reliably. On top of that changes in your work environment can easily pollute LLMs and change quality of results. Are things getting added to your memory or claude.md files that you don't realize? Is your project growing in size and thus claude is performing worse as more context is needed to work with it? etc etc

sigbottle9 days ago
> I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.

The implication is that humans are unreliable and shouldn't be trusted.

Or humans have certain shorthands when they complain on reddit, but their diagnoses are accurate for the specific context? If my AI does something stupid, am I not allowed to call it out? A NS-solving AI is still capable of not satisfying the abstract thing called the user experience. People have intelligent thoughts without compiling to lean.

OK, you say. Then let's get an aggregate benchmark for "intelligence". That doesn't prove that AI didn't flounder a specific use case that the user requested.

Classic moves: Humans are unreliable, converge to some "objective" benchmark that necessarily will quotient out the special cases, etc. Wonder how we'll be solving these issues in the AGI era - well, if you have an AGI that just replicates itself, dominates everybody because it's a machine and humans are soft fleshy creatures, and agrees with itself, fine. But part of the beauty of human experience is the messy part, and providing value is in the messy part.

nomel9 days ago
> and human perception is absolutely horrible at evaluating trends like this

The need to have a mental measure of competence for your fellow man is, most likely, a pre-human skill, probably with a dedicated bit of neurons for it. I think the problem is that those instincts were co-evolved with our fellow man, and, as you say, don't apply at all to a more non-deterministic system that, fundamentally, lacks some logic faculties that even small children have (simple riddle modifications, car wash question, etc).

pixl979 days ago
In other industries of chance we have regulators that ensure compliance and that the providers aren't cheating.

At the end of the day the highest quality of benchmark tells you nothing if the man behind the curtain is constantly changing variables on you. You have no idea if you're really testing the same thing at all. So when you run your test at the top level on their system you're seeing lets say a 30% difference in quality most of the time, you have no idea if you should really only see a 5% difference in quality if you were running a local model with stable settings.

prodigycorp9 days ago
New release of fable and opus 5.5 is pending and Anthropic is reallocating resources. Degradation always happens in transition, it sucks.

Opus 5.5 is being served under opus 5 right now.

gslepak9 days ago
> Opus 5.5 is being served under opus 5 right now.

On what basis are you claiming this?

SequoiaHope9 days ago
Can you elaborate on the mechanism of this degradation? If resources are not available I would expect a request to fail with a message about resources not available. Do they tweak back end model capabilities to maintain service in a degraded state?
pllbnk9 days ago
It shouldn’t be an excuse. They are selling a product and that product should always be within the quality range.
w12969 days ago
Especially with the frequent releases aka version bumps.
pertymcpert9 days ago
Why would reallocating resources make a single inference run worse in quality?
jotato9 days ago
Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.

For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"

I had to tell it to ssh into the server and run journlctl to check it

Anecdotal, I know, but they all seem to be less capable with time.

_edit_ I use the same reasoning level of `medium`

cromka9 days ago
Same exact experience. I worked with both Fable and Sol foe the last two months, daily for several hours, and got used to the very bright, quick thinking, proactive even.

As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par.

Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money.

Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.

reedlaw9 days ago
Same experience with Sonnet on low effort. It used be when I used a "table_name/id" format to reference a db record, it knew exactly how to find it using connected mcp tools. Today it failed 4/4 times (I tried the exact same prompt in 4 separate sessions and each time it replied "I don't have access to [...]"). On medium effort it got it right the first time.
Starlevel0049 days ago
I'm fairly sure it's just luck of the draw if you get put onto a quant'd model or not. I've seen luna xhigh change intelligence fairly drastically on a day to day basis.
pixl979 days ago
Really this is the base problem. You have zero idea where and how your prompt is being executed.

If for example AWS sells you a 2xLarge server there may be some variability in performance but it's going to be averaged out very well.

When it comes to AI services executing your model there is absolutely no information on what and with what settings your model is being executed. Hell, you have no idea if it even is the model you're paying for. Add that models are not deterministic so variability can be pretty large.

This leads to a common set of dynamics that induce cheating behavior in humans. For example, is there a mix of different hardware. Does lessor hardware use different settings? How do you know xhigh is what your prompt ran under. Anthropic has a proven history of running your prompt silently under different models.

This is a huge mess that needs and will be regulated or sued heavily over. Hell, with as many people out there that hate AI it might be easier than one thinks to have a state sue the providers on this and elicit a huge amount of discovery.

ajspig19 days ago
& the nice thing about Hermes (since its open source) is you can be reasonably sure that behavior change is coming from the model and not the harness. (probably)
zahlman9 days ago
FWIW, logged-out ChatGPT claims to be Luna on high. (Unless it's variable for some reason.)
alexjplant9 days ago
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

I wonder what their official explanation for this behavior is.

Wowfunhappy9 days ago
When something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws.

(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)

chrsw9 days ago
I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.

What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

knlam9 days ago
Not true. I can read what Fable output with ease but when it sprout Claudish like Opus 5, I know they are doing something to the model. Yes, you can immediate know the claudish language if you work with opus long enough
espeed9 days ago
They did. More than once...

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...

But it's still happening: https://github.com/anthropics/claude-code/issues/81759

mirashii9 days ago
And here's another great example of how a bunch of people who don't know what's going on throw noise into the system. That post is simply confused: the 1m opus calls are the auto-mode classifier, actual agent calls are still in Fable.
wgd9 days ago
Their exact phrasing IIRC was that they "never intentionally degrade" their models.

This still leaves an absurd amount of wiggle room for arguments like "oh no, our evals show that this quantization has no detectable effect on performance (in the eval distribution) therefore running the quant doesn't degrade quality"

QwenGlazer90009 days ago
Last time they were called out, it was a regression in Claude code itself.

At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.

himata41139 days ago
They are deploying optimizations weekly (if not daily) with various AB tests. They don't manipulate model performance, but they do actively perform tests.

Read the full thread on Hacker News →

Related stories