Behavioral differences from Claude Opus 5 and the prompting and harness patterns that address them: effort calibration, thinking behavior in API integrations and chat, progress updates, unattended and multiagent tasks,…

201 points•Michelangelo11•3 days ago•220 comments•

220 comments

skeledrew3 days ago
All that keeps jumping out at me is how they've set it to refuse giving users thinking tokens and prompts for full reasoning in output. Just drives me further away; I may not stop using Claude completely for now, but I'll be moving even more of my primary workload to Chinese providers. That's where openness and freedom is now at.
asabla2 days ago
> All that keeps jumping out at me is how they've set it to refuse giving users thinking tokens and prompts for full reasoning in output

I keep seeing comments added to code, which reads like reasoning output instead of meaningful words. I see this behavior for both OpenAI and Anthropic models (for several harnesses as well).

But this is a sample of one. And I may be in a situation where I'm more negative to the output from LLMs in general.

dtech2 days ago
Yeah GPT 5.6 models did it a lot and Opus is absolutely awful on this. It's clearly encoding it's thinking/context into the comments. GPT-6 models seem to be better about it.
rajeevk3 days ago
What Chinese models/providers are you using for this? I'm hitting Claude's weekly limits much sooner than I used to with roughly the same workload, so I'm interested in trying alternatives, especially ones with strong coding/agentic performance.
arcanemachiner3 days ago
Get yourself an OpenCode Go subscription and give DeepSeek Flash 4.1 a shot.

A common tactic is to used a big brain model like Opus for planning and reviewing, and a cheaper model for execution.

ffsm82 days ago
> hitting Claude's weekly limits much sooner than I used to with roughly the same workload,

Anthropic had a +50% weekly tokens promotion since April (!) which just ran out last weekend after getting multiple extensions.

I've been feeling that too,and I suspect that's the true reason why they released opus 5.5 at a discount

reachableceo3 days ago
z.ai with zcode. it works around the clock for me off of my redmine queue
surgical_fire3 days ago
I have tried GLM on a subscription, and also DeepSeek and MiMo using API directly. MiMo in particular is extremely cheap.

For regular software development they have been pretty great.

buckle80172 days ago
Don't use opus 5.5 at high. Medium is about as good as 5.0 was at high.
pllbnk3 days ago
Crazy how tables have turned. Life seems surreal since 2020.
baxtr3 days ago
The whole notion of seat-based pricing seems wrong to me as well.
paulirwin2 days ago
Seat-based pricing just includes a certain amount of token usage at a discount for buying in "bulk" (and risking not using all your usage). You can still pay the API token-based rates if you really want to; they won't stop you from doing that.
Jgrubb2 days ago
Can you elaborate?
hannesv3 days ago
What Chinese provider would you use that is on par with Claude code?
arcanemachiner3 days ago
Since Claude Code is a harness that can be made to work with (pretty much?) any model, the answer to the question you have asked is: Claude Code

Non-pedantic answer: I totally agree with you. Opus 5.5 is totally knocking it out of the park IMO.

xandrius3 days ago
Zoo Code is so much better than CC that to me even using similar models I go for CC for simpler things and ZC for larger work.
bluegatty3 days ago
This is a failure of the AI foundries; if we have to use totally different prompting techniques for every model, this wont work.

AI is rapidly saturating it's ability to be useful and these products need to start to mature.

It's not 'fun' to manage 50 different broken MCPs and their variety of ways in which they are broken.

It was 'fun' at the start, now it's just 'broken technology'.

Astra and Opus 5.5 are the 'starting point' for the next era of AI where we expect robust tooling.

ACCount393 days ago
All LLMs understand natural language. All LLMs understand examples. That's honestly more compatibility than you get nearly anywhere, in anything.

The reason why advanced prompting is a moving target is that a lot of prompting is "use extra instructions to compensate for specific ways in which the target LLM is weak or prone to errors". And guess what? LLMs get better over time - obsoleting your advanced prompting.

"Tune a prompt to death for the specific task and the specific model" gets you better performance in the moment, but "trust LLM to be smart" ages a lot more gracefully.

bluegatty3 days ago
"but "trust LLM to be smart" ages a lot more gracefully."

That it doesn't even work now.

The word 'smart' there is actually doing a lot of heavy lifting, it's entirely contextualized.

mpalmer2 days ago

    And guess what? LLMs get better over time - obsoleting your advanced prompting.

It's nowhere near that simple. For instance, models used to be WAY better at writing, until the labs decided that coding ability was a better thing to focus on, and trained successor models accordingly.
TomGarden3 days ago
Agreed.

I'm genuinely worried about all our short term investment in mitigating the failure modes of models that may only be SOTA for a few months.

It's very possible people being 'late' adopting AI may end up with a leg up, not only because they spent more time polishing personal skills during this time, but also because they don't bring all the baggage of 'AI competence' that is becoming irrelevant at breakneck speed.

skybrian3 days ago
Suppose you use LLMs in a more straightforward way, like asking coding agents to make specific changes rather than attempting to build a software factory?

That could be seen either as early adoption that’s overfitted to current capabilities or as late adoption of LLM’s more advanced capabilities.

1980phipsi3 days ago
The people who are adopting LLMs later are also slower to adopt new technology in general. The samples of people who adopt early and adopt late have different characteristics.
pbronez3 days ago
This is “second mover advantage.” There are several dynamics that can make it better to wait and move later. Framed in terms of firms, moving late is advantaged when:

- the product category is long lived

- switching costs are low for buyers

- there are objective standards of quality

- product imitation costs are low

https://insight.kellogg.northwestern.edu/article/the_second_...

Let’s consider those criteria for an individual competing in the labor market with AI. The category should be long lived, AI is here to stay. Switching costs (here, hiring/firing by employers/clients) are low. Objective quality standards fails; technical labor is notoriously difficult to quantify. Imitation costs (can you copy someone else’s good ideas) are moderate but decreasing. That’s where model and tooling improvement shows up.

Based on this analysis, I agree that late movers are well positioned IF the market leaders continue to improve models and tooling to integrate best practices that were previously individual skills.

Early movers should exploit the lack of objective standards. Use your experience with the first generation of tools as marketing to win and retain clients. Continue to invest in soft skills like communication.

maipen3 days ago
What personal skills are you referring to?
saretup3 days ago
> It was 'fun' at the start, now it's just 'broken technology'.

It was even more 'broken' at the start. We overcame some of the issues by 'prompt engineering', which is needed less in the newer, smarter models.

bluegatty3 days ago
Of course - what I mean to say is that we did not perceive it as broken.

The first combustion engine was a miracle. It only becomes 'broken' when we evaluate in some kind of applicable context.

lunchbucket3 days ago
You'll find similar documentation anytime a language or framework or other systems software ships a new major version. It doesn't seem like the way to prompt Opus has changed all that much. Certainly not enough to require a "totally different prompting technique."
rinconrex3 days ago
Counterpoint, the differentiation is maturity. If all models are simply interchangeable commodities, what's the payoff for Anthropic or OpenAI?

Vastly different ways of interacting with each provider is another story, but really we are pretty spoiled here. Slightly different prompting techniques is not really a big deal. If anything it shows the user has some nuance and appreciation for what each model provides.

Fow what it's worth, I am super happy with Opus 5.5. Less verbose than 5 and just gets work done. The progress has been astounding, and if I have to coax it out a bit differently on Opus 5.5 vs Astra 6, I am happy to pay that small price.

prodigycorp3 days ago
Opus 5.5 is a good model, but I've tried to understand the extreme hype about it on social media about Opus' ability to do 2d work, as we got with Astra doing 3d work. In both releases, the models required extensive access to third party apis to generate assets for it, and a lot of the models work was essentially coordinating everything.

There's so many "x generated this in one shot, this is agi" stuff that gives you the impression that you can vibe operate modern models the same way you operated last year's models. There's so much more to it than that. It requires you to put a faith in the leap in the capability of models, one that would've surely been a waste of time in previous models.

Not sure where i'm going with this other than I think most can relate that it's exhausting keeping up with. I cant imagine what it'd be like parenting a kid that went from toddler to puberty in the span of a year and planning for them to go to college the next year. This industry is moving so fast that it's becoming fact that it's the user that's "holding it wrong" every six months.

XenophileJKO3 days ago
Opus 5.5 does NOT need anything other than some javascript/typescript libraries to make very detailed 2d and 3d visualizations. I've spent a week worth of tokens just feeling out what it can do.

The step function change on Opus 5.5 for visual work shocked me.. and I haven't been surprised like this in a long time with LLMs.

EDIT: When I first saw the "P(DOOM)" video and some of the other animations I was VERY skeptical that Opus 5.5 without a lot of tools could make something like that.. until I tried it for myself. It can.. 100%.

Cider99863 days ago
This is the video (I think) in case anyone's wondering: https://www.youtube.com/watch?v=8j-hR4fJywU
prodigycorp3 days ago
It's very good, yes, but I expected it to produce midjourney type results out of the box. That did not happen. The models are definitely granular stuff now though. They must be training off a ton of digital artist stroke data now.

But it other cases, like the music videos, much of the magic is done by access to elevenlabs and suno apis.

Edit: just saw your edit about the pdoom video. Can you share how you prompted it? Would be helpful to know.

collabs3 days ago
I feel like we have different expectations from these frontier models. I don't use Claude code or any agent that has acts to my local machine. I roll up the code and give it the text file that contains all the code. I asked Claude Opus 5.5 max to make me a 2D terminal based racing game with no assets drawings or audio and it exceeded my expectations. Only one failed unit test and that one too it said the test was faulty rather than the code.

I'm still more worried about the malice and any malicious acts by the people at these frontier labs than the models at the frontier labs.

cbg03 days ago
You don't have to run Claude/Codex in auto-approve mode, you can manually approve its interactions with your machine without having to copy your code back and forth between the website and your local files.
alansaber3 days ago
"a lot of the models work was essentially coordinating everything." - I don't see anything wrong with that personally. It's still extremely challenging to build a model harness, and having a model-mediated everything is clearly wishful thinking. It's exhausting to keep up with, but also somewhat exciting, all depends on your perspective of course.
mohamedkoubaa3 days ago
>"x generated this in one shot, this is agi"

Reminds me of the time when you could program a spreadsheet in the 90s and people who didn't know computers would think you were so smart to have invented spreadsheets

bob10293 days ago
> Fourth, if long tool-calling turns still go quiet for longer than you want, have your harness ask for an update

I'm not sure I understand this complexity. In all harnesses I've ever used, tool calls themselves are surfaced to the user as an indication of progress. When the UI/UX around this is engineered well, the user should be able to infer roughly what is going on. Different tools have different ideal presentations. You can't reduce everything to plaintext blobs.

If I absolutely needed intra-turn progress updates, I'd accumulate a separate per-turn transcript and feed it into a cheaper model at deterministic intervals.

adastra223 days ago
Claude code has been hiding tool calls for some months now :(
bob10293 days ago
How does it hide tool calls? I have to run those and return the results.
skerit3 days ago
> In Anthropic's testing, at its default "medium" effort the model matched or beat Claude Opus 5 at "high" effort on such tasks, in fewer steps and with fewer tokens

Opus 5.5 has been amazing, but I'm confused by how this is worded. It "matched or beat" Opus 5? There is no matching. There is only surpassing. By miles. Like Opus 5 was the biggest disappointment of the year. Opus 5.5 is even better than Fable. I do not understand why they're not acknowledging it for the leap that it is?

afro883 days ago
I don't get it either. Ditto for visual design capability. It's so far above Opus 5 and yet the announcement mentioned nothing about it.
philipwhiuk3 days ago
Underneath this means that you have say 50 tests and you grade each of them out of 10, then there was no test it did worse on.

The data doesn't support it being better on every test (sometimes the score will be the same imperfect one, sometimes both will have gotten a perfect score).

Read the full thread on Hacker News →

Related stories