Behavioral differences from Claude Opus 5 and the prompting and harness patterns that address them: effort calibration, thinking behavior in API integrations and chat, progress updates, unattended and multiagent tasks,…
220 comments
I keep seeing comments added to code, which reads like reasoning output instead of meaningful words. I see this behavior for both OpenAI and Anthropic models (for several harnesses as well).
But this is a sample of one. And I may be in a situation where I'm more negative to the output from LLMs in general.
A common tactic is to used a big brain model like Opus for planning and reviewing, and a cheaper model for execution.
Anthropic had a +50% weekly tokens promotion since April (!) which just ran out last weekend after getting multiple extensions.
I've been feeling that too,and I suspect that's the true reason why they released opus 5.5 at a discount
For regular software development they have been pretty great.
Non-pedantic answer: I totally agree with you. Opus 5.5 is totally knocking it out of the park IMO.
AI is rapidly saturating it's ability to be useful and these products need to start to mature.
It's not 'fun' to manage 50 different broken MCPs and their variety of ways in which they are broken.
It was 'fun' at the start, now it's just 'broken technology'.
Astra and Opus 5.5 are the 'starting point' for the next era of AI where we expect robust tooling.
The reason why advanced prompting is a moving target is that a lot of prompting is "use extra instructions to compensate for specific ways in which the target LLM is weak or prone to errors". And guess what? LLMs get better over time - obsoleting your advanced prompting.
"Tune a prompt to death for the specific task and the specific model" gets you better performance in the moment, but "trust LLM to be smart" ages a lot more gracefully.
That it doesn't even work now.
The word 'smart' there is actually doing a lot of heavy lifting, it's entirely contextualized.
And guess what? LLMs get better over time - obsoleting your advanced prompting.
It's nowhere near that simple. For instance, models used to be WAY better at writing, until the labs decided that coding ability was a better thing to focus on, and trained successor models accordingly.I'm genuinely worried about all our short term investment in mitigating the failure modes of models that may only be SOTA for a few months.
It's very possible people being 'late' adopting AI may end up with a leg up, not only because they spent more time polishing personal skills during this time, but also because they don't bring all the baggage of 'AI competence' that is becoming irrelevant at breakneck speed.
That could be seen either as early adoption that’s overfitted to current capabilities or as late adoption of LLM’s more advanced capabilities.
- the product category is long lived
- switching costs are low for buyers
- there are objective standards of quality
- product imitation costs are low
https://insight.kellogg.northwestern.edu/article/the_second_...
Let’s consider those criteria for an individual competing in the labor market with AI. The category should be long lived, AI is here to stay. Switching costs (here, hiring/firing by employers/clients) are low. Objective quality standards fails; technical labor is notoriously difficult to quantify. Imitation costs (can you copy someone else’s good ideas) are moderate but decreasing. That’s where model and tooling improvement shows up.
Based on this analysis, I agree that late movers are well positioned IF the market leaders continue to improve models and tooling to integrate best practices that were previously individual skills.
Early movers should exploit the lack of objective standards. Use your experience with the first generation of tools as marketing to win and retain clients. Continue to invest in soft skills like communication.
It was even more 'broken' at the start. We overcame some of the issues by 'prompt engineering', which is needed less in the newer, smarter models.
The first combustion engine was a miracle. It only becomes 'broken' when we evaluate in some kind of applicable context.
Vastly different ways of interacting with each provider is another story, but really we are pretty spoiled here. Slightly different prompting techniques is not really a big deal. If anything it shows the user has some nuance and appreciation for what each model provides.
Fow what it's worth, I am super happy with Opus 5.5. Less verbose than 5 and just gets work done. The progress has been astounding, and if I have to coax it out a bit differently on Opus 5.5 vs Astra 6, I am happy to pay that small price.
There's so many "x generated this in one shot, this is agi" stuff that gives you the impression that you can vibe operate modern models the same way you operated last year's models. There's so much more to it than that. It requires you to put a faith in the leap in the capability of models, one that would've surely been a waste of time in previous models.
Not sure where i'm going with this other than I think most can relate that it's exhausting keeping up with. I cant imagine what it'd be like parenting a kid that went from toddler to puberty in the span of a year and planning for them to go to college the next year. This industry is moving so fast that it's becoming fact that it's the user that's "holding it wrong" every six months.
The step function change on Opus 5.5 for visual work shocked me.. and I haven't been surprised like this in a long time with LLMs.
EDIT: When I first saw the "P(DOOM)" video and some of the other animations I was VERY skeptical that Opus 5.5 without a lot of tools could make something like that.. until I tried it for myself. It can.. 100%.
But it other cases, like the music videos, much of the magic is done by access to elevenlabs and suno apis.
Edit: just saw your edit about the pdoom video. Can you share how you prompted it? Would be helpful to know.
I'm still more worried about the malice and any malicious acts by the people at these frontier labs than the models at the frontier labs.
Reminds me of the time when you could program a spreadsheet in the 90s and people who didn't know computers would think you were so smart to have invented spreadsheets
I'm not sure I understand this complexity. In all harnesses I've ever used, tool calls themselves are surfaced to the user as an indication of progress. When the UI/UX around this is engineered well, the user should be able to infer roughly what is going on. Different tools have different ideal presentations. You can't reduce everything to plaintext blobs.
If I absolutely needed intra-turn progress updates, I'd accumulate a separate per-turn transcript and feed it into a cheaper model at deterministic intervals.
Opus 5.5 has been amazing, but I'm confused by how this is worded. It "matched or beat" Opus 5? There is no matching. There is only surpassing. By miles. Like Opus 5 was the biggest disappointment of the year. Opus 5.5 is even better than Fable. I do not understand why they're not acknowledging it for the leap that it is?
The data doesn't support it being better on every test (sometimes the score will be the same imperfect one, sometimes both will have gotten a perfect score).
Read the full thread on Hacker News →
Related stories
- Hacker News · 2 points · 7 days ago
- DEV Community · 0 points · about 8 hours ago
- The Verge · 0 points · 8 days ago
- Hacker News · 4 points · 12 days ago
- Claude Opus 5.5anthropic.comHacker News · 1788 points · 8 days ago
- Claude Opus 5.5anthropic.comHacker News · 273 points · 8 days ago