Strands harness is a fully assembled, customizable, state-of-the-art agent you run locally or deploy anywhere.

148 points•zuckerborg0101•8 days ago•96 comments•

96 comments

johnmlussier7 days ago
I am increasingly hesitant to use non-native harnesses - model providers are now starting to train their agents for use within the harness. An eval like terminal bench can only capture so much data. I don't want to have to assess each harness every model release to make sure it's working as well as it can.
verdverm7 days ago
Interestingly, MiMo was trained across multiple harnesses and it improved capabilities

https://mimo.mi.com/docs/en-US/news/latest/v2-6

SwellJoe7 days ago
MiMo _needed_ it. MiMo 2.5 Pro really struggled when given more tools, causing it to lose other capabilities. e.g. in security benchmarks I was doing, most of the medium-to-large models got better or stayed roughly the same at finding security vulnerabilities when given more tools (e.g. treesitter, semgrep, a full bash/python environment, etc.) vs. when only given the ability to read the files in the repo. But, MiMo got notably worse. It seemed to get confused by all the options. MiMo 2.5 Pro just reading files is an excellent security bug finder, at the pareto frontier at the time (cheapest option to find as many bugs as Opus 4.8, which was current at the time), but adding tools cratered it down to small model territory.

I haven't tested the new version, and I still have a bunch of tokens on my token plan, so I might give 2.6 a go in some similar tests to see if it still has the analysis paralysis problem of 2.5.

avaer7 days ago
This matters less as models get better and everyone settles on the same overall harness architectures. The model matters more than the harness anyway.

The bigger issue is that the use cases and harnesses for models is infinite, which is hard to compress into benchmark numbers that actually apply to you.

Everyone is benchmaxxing, desperate to sell, and almost nobody except the labs is doing actual science on the results, so harnesses tend to be chosen on voodoo and hunches, like which company made it. There isn't necessarily a good alternative though, bearing the cost of being a harness researcher is probably not many people's goal.

CharlieDigital7 days ago

    > The model matters more than the harness anyway
    > 
    > Everyone is benchmaxxing
    > 
    > ...harnesses tend to be chosen on voodoo and hunches...
I get what you're saying, but their graphic on performance here uses the exact same model with different harnesses and definitively shows that there is a significant difference in both accuracy and cost. The whole point of their technical implementation and design decision here is to highlight that it's not "voodoo and hunches", but observable data.

Fable 5 on Claude Code scored 61.8% at a cost of $248.05 while Fable 5 on OpenCode beat it at 66.3% at $73.42. The same model, the same benchmark; only the harness is different with a ~5 point difference in accuracy while costing significantly less. So if we are to believe the author and these results are repeatable, then it would seem that the harness matters.

The point of this framing here is specifically to address 1) benchmaxxing by using the same model, 2) NOT choose a harness on "voodoo and hunches" by using actual data to back the assertions. Your comment feels misguided and completely hand waves the actual data points here.

altcognito7 days ago
Agreed, especially since the more frontier models are able to accomplish in a vacuum, the more people will trust them. That being said, tool use is still really important for pulling in the right information.
stogot7 days ago
I’m the opposite. I want one open source harness to rule them all

Cost efficiency is a plus

sanderjd7 days ago
Right! I had the exact opposite view when reading this comment. The good timeline is where one of (or perhaps a small number of) the open source harnesses becomes so dominant that the models compete to have the model that is the best trained to work with that harness. (I'm hoping this would be Pi, because it's my favorite, but mostly I just want it to be some model-agnostic open source harness that wins.)
crossroadsguy7 days ago
I mostly use GLM. But I will not touch ZAI's harness with a 10 mile long pole no matter how efficiently they couple it with GLM.
UncleOxidant7 days ago
Is your main concern privacy? I've been using the zcode of and on for a few months now and it's definitely improved over what it was back in the spring. I also use DeepSeek's harness. Not sure which I prefer at this point. Used to be I preferred DSH, but zcode has some features I like over DSH.
jorgeleo7 days ago
This is the very reason why I avoid to use native harness. They are optimized for the economic benefit of the provider, not for mine. Love Pi because it does a good job managing the context, open code meh, claude code nope.
theturtletalks8 days ago
Why is Pi not in the benchmarks? Deepseek beats Strands and its built on Pi so that’s all I needed to know.
cobolcomesback7 days ago
Base Pi doesn’t seem like an apt comparison here. Strands comes with MCP servers, subagents, and web fetch tools built in. To get those on Pi, you have to add addons, and then you get into the space of “which addons should be added to make it an apples-to-apples comparison? Which MCP addon do I use, the fastest one, or the most popular one?”. My guess is they chose OMP because it’s the defacto standard for answering those “which addons should be considered out of the box” questions for Pi.
theturtletalks7 days ago
I figured it was lack of MCP support. I would just include a note why it was excluded. Have you checked out FX? It seems to be like Pi but will more things out of the box like MCP.
miroljub7 days ago
If "pi install npm:pi-mcp-adapter" is preventing them from including pi in benchmarks, then they might not be competent to run trustworthy benchmarks.
leodavi7 days ago
Deepseek harness is not actually built on Pi harness. It's an independent project.
theturtletalks7 days ago
This is true, but the team said on Twitter that they were heavily inspired by Pi and this whole team used it extensively:

https://x.com/tianyi/status/2088306143772946499

techscruggs7 days ago
Deepseek was cheaper, but also less accurate. "Beats" isn't a fair assessment.
theturtletalks7 days ago
That’s fair, but Strands is advertising their harness needing way less tokens which does relate to cost.

In that same vein, Pi is less bloated then Deepseek and Oh My Pi, which are built on top of Pi. Isn’t it dubious to leave it out?

crossroadsguy7 days ago
OhMyPi is essentially Claude like bloated and Pi is so barebones that many will just get frustrated trying to add bricks after bricks to make it usable for complex workflows - so I don't really see any point of it being in such benchmarks. Pi should have added option to easily add "packs" (and then also tweak tehm) sort of things. When I tried it I actually thought that's what OMP will be, but it was not. It was a different harness altogether in a way and not a nice way.
scuppernong7 days ago
I wish we could use just the frontend of omp harness-agnostically, like the provider switching, stats dashboard, subscription pooling etc. are excellent features, but it doesn't seem inconceivable that we can have all that with a swappable "actual" harness. I've tried to make omp more pi-like by lazy loading most of the tools rather than dumping them into system prompt. haven't benchmarked this yet.
jsw977 days ago
Came here to say this. They have oh-my-pi in the benchmark but not pi, but those are very different animals. pi is lightweight out of the box so has very little start-time overhead. (And will not spin up agents like crazy.) pi might do worse if those things are actually important for solving the problem, but it certainly has a shot at being most efficient.
strandstan7 days ago
hi Albert from the Strands team. We can look into a Pi run. Our researcher is working on a deep dive for our benchmarks, so we got resources to test against other harnesses
seizethecheese7 days ago
> With Fable 5, Strands harness cost 77% less than Claude Code and scored higher on Terminal Bench 2.1.

Terminal Bench 2.1 is saturated. Many token saving techniques would save money and score basically the same running Fable 5 against Terminal Bench 2.1. (They claim a better score but don’t say how much better. I’d bet my favorite hat that it’s not statistically significant.)

This is at least the fourth time I’ve seen a project hit front page with a “save money with same score on saturated benchmark” claim.

strandstan7 days ago
The scores are in blog post's bar chart. For Terminal Bench 2.1, Strands harness (Fable 5) scored 69.7 while Claude Code (Fable 5) scored 61.8. This is on high effort.

I hear you tho about saturation. We're working on a follow-up deep dive post with more harnesses, so could look into Terminal Bench 4.0?

seizethecheese7 days ago
Then I'm really confused. Terminal Bench 2.1 scores on Artificial Analysis are like 80-90%. https://artificialanalysis.ai/evaluations/terminalbench-2-1
Oras7 days ago
I remember reading about strands SDK and it looked great in terms how everything is an event that you can extend, so this harness feels quite about right.

However, for this kind of customisation, Pi is actually quite great. One of the most things I love about Pi is ability to ask it to create an extension and it does it quite well as it’s part of their docs. Also ability to customise the system prompt to avoid the clutter that Claude Code add (around 20k system prompt that mostly had nothing to do with the code).

The demo was showing something I have created for my Pi setup, which is asking me in each new session which skills and MCP I would to enable for the session. This works quite well if you have multiple projects where you don’t need all skills but just a small subset

samusiam7 days ago
> But the moment you build your own agent, you’re on your own. It’s tricky wiring up the right primitives just well enough to match that “it just worked” feeling.

It's wild to me to claim that it's tricky to customize one of these harnesses and for that to be the entire justification for an entirely different harness.

It's really not that hard. If you want to reduce costs then all you need to do is practice delegation: instead of using the strong model, all the time to do everything, instead, you have the stronger model delegate well-defined tasks to a weaker model. Patterns like these are really easy to wire up.

seizethecheese7 days ago
Yes, it’s a wild claim. I built a harness for a random side project without even thinking hard about it. The harness was that from the hardest part of the project.

Is this corporate confabulation?

Read the full thread on Hacker News →

Related stories