harness

63 stories and discussions about harness, aggregated from every source we track.

1.

Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the…

221 points•wek•13 days ago•59 comments•
2.

We’re sharing Unreal Agent — an agent harness that delivers up to 40% cost savings compared to Codex on production workloads and coding/science benchmarks, without any negative performance impact.

217 points•trollied•8 days ago•114 comments•
3.

Strands harness is a fully assembled, customizable, state-of-the-art agent you run locally or deploy anywhere.

129 points•zuckerborg0101•8 days ago•89 comments•
4.

The Harness of Harnesses • built for RSI: a trusted, persistent, self-evolving multi-agent ecosystem for all-domain collaboration. - EverMind-AI/Raven

53 points•cyfyifanchen•1 day ago•48 comments•
5.

How today's models perform on computer use benchmarks, compared on accuracy, cost per task, and wall-clock speed.

20 points•MiguelG719•6 days ago•15 comments•
6.

This is a submission for the MLH x DEV Writing Challenge What I Built AI products are...

10 points•rajan_mishra_a9f78ad216b4•4 days ago•1 comment
7.

A serverless harness for agents to create custom agents and deploy them as tools, MCPs, or bots.

8 points•ozankabak•5 days ago•4 comments•
8.

It means giving the agent enough context to do the right thing For this week's Agent Factory...

8 points•annthurium•16 days ago•4 comments
9.

Event-driven, harness-neutral API and CLI for live coding-agent sessions - markwylde/all-your-agents

6 points•turblety•9 days ago•4 comments•
10.

Local-first agent harness (Windows + Ollama) for 2–9B models. Same 2B model: 0.017 → 0.821 across four agent harnesses — 48×. 288-cell benchmark, every cell public. One-click zero-outbound mode. 本地...

5 points•Gustor•about 12 hours ago•0 comments•
11.
12.
4 points•grigio•8 days ago•10 comments•
13.

As models improve they absorb the harness: planning, tool use, retries, and self-checking move into the weights. They cannot absorb the layer that ends the regress of enforcement. An operating system for agents may…

3 points•kendallgclark•about 10 hours ago•0 comments•
14.
3 points•jinhongyii•1 day ago•0 comments•
15.

The only harness you need for coding. Relay keeps what every run learns, re-tests every task, and runs on your machine with your model.

3 points•rsathwik07•1 day ago•0 comments•
16.

Manage a team of AI agents to run your business. Org charts, budgets, governance, and goals — all in one deployment.

3 points•FloatArtifact•7 days ago•0 comments•
17.

Anthropic and OpenAI both shipped a new model yesterday. We ran them through our hard cases overnight. Here is what changed, and why partforge is still on Gemini.

3 points•Mazer23•7 days ago•0 comments•
18.

Strands harness is a fully assembled, customizable, state-of-the-art agent you run locally or deploy anywhere.

3 points•fourfire•8 days ago•0 comments•
19.

Measurement harness for proxy providers, browser engines and scraping targets. Every claim carries the run it came from. - nodemaven/proxy-benchmark

2 points•pia-nm•about 16 hours ago•0 comments•
20.

Open source agent harness, built from the ground up for cloud-native production workloads on infrastructure you control. Run the same provider-agnostic loop locally, remotely, or on Kubernetes, wit...

2 points•kantord•1 day ago•0 comments•
21.

A cybersecurity harness for full-stack LLM-driven penetration testing. Find and fix vulnerabilities autonomously, 24/7. [RESEARCH PREVIEW] - 0sec-labs/0

2 points•soltanov•1 day ago•0 comments•
22.

Multi-agent harness that runs Claude Code and Codex together as one system - mvschwarz/openrig

2 points•handfuloflight•3 days ago•0 comments•
23.

Zhening Li, Omar Khattab, Armando Solar-Lezama and colleagues at MIT CSAIL build JAZ, an agent framework whose only primitive is an LLM-backed `invoke` function

2 points•omarsar•3 days ago•0 comments•
24.

Claude spent 119,000 tokens to produce a 107-token program that didn't work, in a programming language that nobody uses? Or that the code you make with Claude is about 2% of your token spend? If so, read this.

2 points•BenediktHolm•4 days ago•0 comments•
25.

Keep it simple, stupid. A performant agent harness inspired off Pi. - racetozero/kiss

2 points•racetozero•5 days ago•2 comments•
26.

Isn't agent memory just a bunch of markdown files with an MCP? Sometimes! But there's real nuance in how and when you store and retrieve. An interactive tour of every way agents can remember.

2 points•shibel•6 days ago•0 comments•
27.

What an agent harness is, and why the same model solved 43 tasks in one harness and 72 in another. With Claude Code, Codex and six more coding agents compared.

2 points•arifulislamat•6 days ago•3 comments
29.

Contribute to ai-cad-labs/ai-cad development by creating an account on GitHub.

2 points•jbm•7 days ago•0 comments•
30.

Deterministic agent harness for Temporal (in Rust) - smartcomputer-ai/lightspeed

2 points•handfuloflight•8 days ago•0 comments•
31.

An AI coding agent that doesn't read - it queries. 391 of 500 real GitHub issues resolved at 9.5 cents per fix.

2 points•tweedler290•8 days ago•1 comment•
32.

A Windows-local, fail-closed supervisor for long-running AI work loops. Plant clock ≠ chat — chat is a mouth; the work clock is not the chat. - denisrigsby/Aetheria

2 points•denisrigsby•8 days ago•0 comments•
33.

Measurement harness for proxy providers, browser engines and scraping targets. Every claim carries the run it came from. - nodemaven/proxy-benchmark

2 points•pia-nm•10 days ago•0 comments•
34.

A decision layer embedded into the harness tool-selection loop

1 points•raezil•about 17 hours ago•0 comments•
35.

After a long hiatus, the problem, which was likely inspired by juggling, has finally been resolved by a group of young mathematicians.

1 points•pykello•about 18 hours ago•0 comments•
36.

After a long hiatus, the problem, which was likely inspired by juggling, has finally been resolved by a group of young mathematicians.

1 points•tzury•1 day ago•0 comments•
37.

AI agents guess, and they guess fast. Linters, type checkers, the compiler, formatters, and a real test bed are what turn those guesses into verified changes. The tighter that loop, the more of the build an agent can…

1 points•bostonaholic•1 day ago•0 comments•
38.

Talk to Claude Code while it works. A full-duplex voice companion powered by GPT Live 1. - CakeCrusher/full_duplex_code

1 points•SebastianSosa•1 day ago•0 comments•
39.

We built a loop that lets software mutate and improve itself, then applied it to a coding-agent harness. Five generations in, it's more accurate and cheaper to run.

1 points•CharlieDigital•1 day ago•0 comments•
40.

After a long hiatus, the problem, which was likely inspired by juggling, has finally been resolved by a group of young mathematicians.

1 points•pseudolus•1 day ago•0 comments•
41.

Manus 2.0 is here, with the new Cascade agent harness, Cloud Computer and Automations, Manus Studio for professional creation, and Cue, a new app for personal agents.

1 points•shenli3514•2 days ago•2 comments•
42.

After a long hiatus, the problem, which was likely inspired by juggling, has finally been resolved by a group of young mathematicians.

1 points•ibobev•2 days ago•0 comments•
43.

Chock is a sandbox-first AI coding harness. It runs a coding agent inside your operating system's own sandbox, in a throwaway copy of your project, and writes every turn to a session log you can read afterwards.

1 points•rosscomputerguy•4 days ago•0 comments•
44.

Some short musings on the shape of language models, e.g. what it means to design a language model around a harness, and not the other way around.

1 points•emersonmacro•4 days ago•0 comments•
45.

Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task…

1 points•acossta•4 days ago•0 comments•
46.

The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime. - dream-num/univer

1 points•qwbfsa•5 days ago•0 comments•
47.

Everyone's defining the ideal coding agent. The company-brain harness people actually want knows the company, clears urgent work, stays multiplayer — and stays open to the agents they already use.

1 points•ferran9908•5 days ago•0 comments•
48.

A full agentic loop delivered to your phone. Contribute to nev3rfail/harness.apk development by creating an account on GitHub.

1 points•nev3rfail•6 days ago•0 comments•
49.
1 points•grigio•7 days ago•0 comments•
50.

NetHack is hard to hack. But is it really? Let's find out together.

1 points•vokneruk•7 days ago•2 comments•
51.

what belongs in a research harness when the models keep improving, and how do we get the expertise into it?

1 points•galsapir•7 days ago•0 comments•
52.

A programmable control plane where every coding harness and your local dev stack can work together.

1 points•ibobev•7 days ago•0 comments•
53.

Build an agent harness on the Pi SDK that uses Jev to pick models, block risky tool calls, and check its own answers.

1 points•omarsar•8 days ago•0 comments•
54.

Async-first agent harness. Contribute to unreallabsai/unreal-agent development by creating an account on GitHub.

1 points•Bluestein•8 days ago•0 comments•
55.

Introducing Oh My Quant: a research-native agent harness for experiments, durable evidence, and traceable conclusions.

1 points•RuiWang0811•8 days ago•1 comment•
56.

I ran five different coding agents at the same local model, on the same task, with the same frozen test suite — and then I counted why they failed. The answer wasn’t subtle. About 90% of the …

1 points•gherlein•8 days ago•0 comments•
57.

We pointed our autoresearch loop at Jev, a classifier that can't be fine-tuned. It ended up making half as many mistakes as a reasoning model, for a seventh of the cost and a thirtieth of the latency.

1 points•scosman•9 days ago•0 comments•
58.

Orcrist is a desktop coding agent whose harness is written fresh for every task - simone20a/Orcrist

1 points•simone20a•9 days ago•0 comments•
59.

Austin's go-to studio for innovative mobile apps and games. Build impactful apps for artists and entrepreneurs with our AI-powered solutions.

1 points•bishopz•9 days ago•0 comments•
60.

Strands harness is a fully assembled, customizable, state-of-the-art agent you run locally or deploy anywhere.

1 points•ot•9 days ago•0 comments•

Related topics