harness
63 stories and discussions about harness, aggregated from every source we track.
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the…
We’re sharing Unreal Agent — an agent harness that delivers up to 40% cost savings compared to Codex on production workloads and coding/science benchmarks, without any negative performance impact.
Strands harness is a fully assembled, customizable, state-of-the-art agent you run locally or deploy anywhere.
The Harness of Harnesses • built for RSI: a trusted, persistent, self-evolving multi-agent ecosystem for all-domain collaboration. - EverMind-AI/Raven
How today's models perform on computer use benchmarks, compared on accuracy, cost per task, and wall-clock speed.
This is a submission for the MLH x DEV Writing Challenge What I Built AI products are...
A serverless harness for agents to create custom agents and deploy them as tools, MCPs, or bots.
It means giving the agent enough context to do the right thing For this week's Agent Factory...
Event-driven, harness-neutral API and CLI for live coding-agent sessions - markwylde/all-your-agents
Local-first agent harness (Windows + Ollama) for 2–9B models. Same 2B model: 0.017 → 0.821 across four agent harnesses — 48×. 288-cell benchmark, every cell public. One-click zero-outbound mode. 本地...
As models improve they absorb the harness: planning, tool use, retries, and self-checking move into the weights. They cannot absorb the layer that ends the regress of enforcement. An operating system for agents may…
The only harness you need for coding. Relay keeps what every run learns, re-tests every task, and runs on your machine with your model.
Manage a team of AI agents to run your business. Org charts, budgets, governance, and goals — all in one deployment.
Anthropic and OpenAI both shipped a new model yesterday. We ran them through our hard cases overnight. Here is what changed, and why partforge is still on Gemini.
Strands harness is a fully assembled, customizable, state-of-the-art agent you run locally or deploy anywhere.
Measurement harness for proxy providers, browser engines and scraping targets. Every claim carries the run it came from. - nodemaven/proxy-benchmark
Open source agent harness, built from the ground up for cloud-native production workloads on infrastructure you control. Run the same provider-agnostic loop locally, remotely, or on Kubernetes, wit...
A cybersecurity harness for full-stack LLM-driven penetration testing. Find and fix vulnerabilities autonomously, 24/7. [RESEARCH PREVIEW] - 0sec-labs/0
Multi-agent harness that runs Claude Code and Codex together as one system - mvschwarz/openrig
Zhening Li, Omar Khattab, Armando Solar-Lezama and colleagues at MIT CSAIL build JAZ, an agent framework whose only primitive is an LLM-backed `invoke` function
Claude spent 119,000 tokens to produce a 107-token program that didn't work, in a programming language that nobody uses? Or that the code you make with Claude is about 2% of your token spend? If so, read this.
Keep it simple, stupid. A performant agent harness inspired off Pi. - racetozero/kiss
Isn't agent memory just a bunch of markdown files with an MCP? Sometimes! But there's real nuance in how and when you store and retrieve. An interactive tour of every way agents can remember.
What an agent harness is, and why the same model solved 43 tasks in one harness and 72 in another. With Claude Code, Codex and six more coding agents compared.
Contribute to ai-cad-labs/ai-cad development by creating an account on GitHub.
Deterministic agent harness for Temporal (in Rust) - smartcomputer-ai/lightspeed
An AI coding agent that doesn't read - it queries. 391 of 500 real GitHub issues resolved at 9.5 cents per fix.
A Windows-local, fail-closed supervisor for long-running AI work loops. Plant clock ≠ chat — chat is a mouth; the work clock is not the chat. - denisrigsby/Aetheria
Measurement harness for proxy providers, browser engines and scraping targets. Every claim carries the run it came from. - nodemaven/proxy-benchmark
A decision layer embedded into the harness tool-selection loop
After a long hiatus, the problem, which was likely inspired by juggling, has finally been resolved by a group of young mathematicians.
After a long hiatus, the problem, which was likely inspired by juggling, has finally been resolved by a group of young mathematicians.
AI agents guess, and they guess fast. Linters, type checkers, the compiler, formatters, and a real test bed are what turn those guesses into verified changes. The tighter that loop, the more of the build an agent can…
Talk to Claude Code while it works. A full-duplex voice companion powered by GPT Live 1. - CakeCrusher/full_duplex_code
We built a loop that lets software mutate and improve itself, then applied it to a coding-agent harness. Five generations in, it's more accurate and cheaper to run.
After a long hiatus, the problem, which was likely inspired by juggling, has finally been resolved by a group of young mathematicians.
Manus 2.0 is here, with the new Cascade agent harness, Cloud Computer and Automations, Manus Studio for professional creation, and Cue, a new app for personal agents.
After a long hiatus, the problem, which was likely inspired by juggling, has finally been resolved by a group of young mathematicians.
Chock is a sandbox-first AI coding harness. It runs a coding agent inside your operating system's own sandbox, in a throwaway copy of your project, and writes every turn to a session log you can read afterwards.
Some short musings on the shape of language models, e.g. what it means to design a language model around a harness, and not the other way around.
Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task…
The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime. - dream-num/univer
Everyone's defining the ideal coding agent. The company-brain harness people actually want knows the company, clears urgent work, stays multiplayer — and stays open to the agents they already use.
A full agentic loop delivered to your phone. Contribute to nev3rfail/harness.apk development by creating an account on GitHub.
NetHack is hard to hack. But is it really? Let's find out together.
what belongs in a research harness when the models keep improving, and how do we get the expertise into it?
A programmable control plane where every coding harness and your local dev stack can work together.
Build an agent harness on the Pi SDK that uses Jev to pick models, block risky tool calls, and check its own answers.
Async-first agent harness. Contribute to unreallabsai/unreal-agent development by creating an account on GitHub.
Introducing Oh My Quant: a research-native agent harness for experiments, durable evidence, and traceable conclusions.
I ran five different coding agents at the same local model, on the same task, with the same frozen test suite — and then I counted why they failed. The answer wasn’t subtle. About 90% of the …
We pointed our autoresearch loop at Jev, a classifier that can't be fine-tuned. It ended up making half as many mistakes as a reasoning model, for a seventh of the cost and a thirtieth of the latency.
Orcrist is a desktop coding agent whose harness is written fresh for every task - simone20a/Orcrist
Austin's go-to studio for innovative mobile apps and games. Build impactful apps for artists and entrepreneurs with our AI-powered solutions.
Strands harness is a fully assembled, customizable, state-of-the-art agent you run locally or deploy anywhere.