Strands harness is a fully assembled, customizable, state-of-the-art agent you run locally or deploy anywhere.
96 comments
I haven't tested the new version, and I still have a bunch of tokens on my token plan, so I might give 2.6 a go in some similar tests to see if it still has the analysis paralysis problem of 2.5.
The bigger issue is that the use cases and harnesses for models is infinite, which is hard to compress into benchmark numbers that actually apply to you.
Everyone is benchmaxxing, desperate to sell, and almost nobody except the labs is doing actual science on the results, so harnesses tend to be chosen on voodoo and hunches, like which company made it. There isn't necessarily a good alternative though, bearing the cost of being a harness researcher is probably not many people's goal.
> The model matters more than the harness anyway
>
> Everyone is benchmaxxing
>
> ...harnesses tend to be chosen on voodoo and hunches...
I get what you're saying, but their graphic on performance here uses the exact same model with different harnesses and definitively shows that there is a significant difference in both accuracy and cost. The whole point of their technical implementation and design decision here is to highlight that it's not "voodoo and hunches", but observable data.Fable 5 on Claude Code scored 61.8% at a cost of $248.05 while Fable 5 on OpenCode beat it at 66.3% at $73.42. The same model, the same benchmark; only the harness is different with a ~5 point difference in accuracy while costing significantly less. So if we are to believe the author and these results are repeatable, then it would seem that the harness matters.
The point of this framing here is specifically to address 1) benchmaxxing by using the same model, 2) NOT choose a harness on "voodoo and hunches" by using actual data to back the assertions. Your comment feels misguided and completely hand waves the actual data points here.
Cost efficiency is a plus
In that same vein, Pi is less bloated then Deepseek and Oh My Pi, which are built on top of Pi. Isn’t it dubious to leave it out?
Terminal Bench 2.1 is saturated. Many token saving techniques would save money and score basically the same running Fable 5 against Terminal Bench 2.1. (They claim a better score but don’t say how much better. I’d bet my favorite hat that it’s not statistically significant.)
This is at least the fourth time I’ve seen a project hit front page with a “save money with same score on saturated benchmark” claim.
I hear you tho about saturation. We're working on a follow-up deep dive post with more harnesses, so could look into Terminal Bench 4.0?
However, for this kind of customisation, Pi is actually quite great. One of the most things I love about Pi is ability to ask it to create an extension and it does it quite well as it’s part of their docs. Also ability to customise the system prompt to avoid the clutter that Claude Code add (around 20k system prompt that mostly had nothing to do with the code).
The demo was showing something I have created for my Pi setup, which is asking me in each new session which skills and MCP I would to enable for the session. This works quite well if you have multiple projects where you don’t need all skills but just a small subset
It's wild to me to claim that it's tricky to customize one of these harnesses and for that to be the entire justification for an entirely different harness.
It's really not that hard. If you want to reduce costs then all you need to do is practice delegation: instead of using the strong model, all the time to do everything, instead, you have the stronger model delegate well-defined tasks to a weaker model. Patterns like these are really easy to wire up.
Is this corporate confabulation?
Read the full thread on Hacker News →
Related stories
- Strands harness: frontier performance with 28% lower token coststrandsagents.comHacker News · 3 points · 8 days ago
- Strands harness: frontier performance with 28% lower token coststrandsagents.comHacker News · 1 points · 9 days ago
- Hacker News · 1 points · 10 days ago
- State of emergency! The four ways your state might be wrongnote89.github.ioLobsters · 9 points · over 4 years ago
- DEV Community · 15 points · 12 days ago
- DEV Community · 8 points · 20 days ago