Every byte verified. Your coding assistant plans, sandboxed workers build, and nomArmy checks every change before it's committed. - rayson-tech/nomarmy

3 points•RaysonTech•4 days ago•1 comment•

1 comment

RaysonTech4 days ago
I’ve been using multiple coding agents across real projects and kept running into a problem that bothered me more than model selection or orchestration:

Coding agents are pretty good at telling you they finished something they didn’t actually finish.

So I built Nom Army and recently open-sourced it.

The architecture is a General that keeps intent and acceptance authority while bounded workers implement pieces of the plan. Workers can be Claude Code, Codex, Grok, Muse, local models, etc. Different models are useful for different work, so I’d rather delegate a bounded job to the right worker than keep swapping models inside one context.

The part I care about most, though, is being able to walk away.

Nom Army treats a worker’s response as a claim, not evidence. Workers operate in isolated Git worktrees/sandboxes. When one reports it’s finished, Nom Army independently examines repository state, reruns verification, scans for secrets, and can revert the production change and rerun the relevant test to establish whether that test actually proves the change.

I now have some early numbers from real usage.

Since Sep 25, one repository has put 42 jobs through Nom Army: 31 implementation jobs and 11 scouts, plus 3 verify-only runs. Workers included Codex/GPT, Grok, and a local 20B model.

Of 21 implementation jobs that reported “done”, 5 (24%) did not survive Nom Army’s evidence checks:

* 3 failed independent verification despite reporting that tests passed. * 2 passed verification, but the relevant tests still passed after Nom Army reverted the production change.

That’s exactly the failure mode I built this for.

But here’s the more interesting result: verification isn’t truth either.

Of the 15 jobs that passed both checks, I still found real defects in 5 during integration/review. These included deployment-bundle failures not exercised locally, a backward-incompatible contract change, an incorrect security classification, and a cross-tenant authorization problem.

You can independently verify the wrong things really well.

That pushed me further toward repository-aware verification. Nom Army has harnesses for Node, Python, Go, Rust, Playwright/Chromium and isolated fake services. Validation is pluggable, so additional systems can independently evaluate a job.

I’m increasingly convinced the worker, orchestrator, test harness and validator shouldn’t necessarily be the same system or even trust one another.

The numbers haven’t all been flattering. There were timeouts, oversized briefs, infrastructure failures and truncated scout reports. Earlier experiments showed that small, already-diagnosed tasks can consume substantially more General tokens when delegated than when the General fixes them directly. More local workers on the same hardware didn’t automatically improve throughput either.

In this run, the median implementation worker ran for 8.7 minutes (p90 16.6). Nom Army committed changes from 16 jobs totaling +4,561/-967 lines across 63 files, including 15 new test files.

I’m less interested in claiming this is token-efficient than attention-efficient. I’ve been using it across several projects concurrently, and the useful part is delegating work and coming back to evidence rather than continuously babysitting coding sessions.

I’m publishing the failures and findings because I don’t think multi-agent coding magically makes software cheaper or correct. It mostly creates another distributed system with some unusually confident components.

Nom Army is Apache-2.0 and still alpha. Real usage has already exposed plenty of problems in Nom Army itself that I’ve fixed in the last two releases.

I’d particularly appreciate criticism around the trust model, sandbox boundaries, revert methodology, harness architecture, and what evidence should exist before an autonomous coding job deserves to be accepted.

Worker output is a claim; repository state is evidence.

Have at it.

Read the full thread on Hacker News →

Related stories