Spend tokens on judgment, not typing. An MCP-native autonomous SDLC pipeline: frontier models plan and review, local models implement — engineering discipline on a $20/month budget. - motock/fagan

2 points•motock•6 days ago•1 comment•

1 comment

motock6 days ago
The intent is to spend tokens on judgment, not typing: enterprise-grade engineering discipline on a $20/month budget. Frontier models are expensive and good at judgment; open-weight models are cheap or free and good enough at writing code once the work is well specified.

Fagan is an MCP server that runs a small SDLC pipeline on that split. Gates sit between every step: tests written before the code, acceptance-oracle grading, an independent review, a merge gate that re-runs the suite on the rebased branch, and a policy layer that stops for a human on anything irreversible. Every role (decompose, plan, implement, review, adjudicate) has its own provider and model: a frontier model, Ollama, MLX, LM Studio, or LiteLLM. One registry sets the defaults, and a plan or story can override them.

It's dogfooded: 802 of the repo's 968 merged PRs went through its own pipeline. Of the 467 with surviving manifests, 93% were implemented by open-weight models: glm-5.3-flash 72%, gpt-oss-20b on-device 11%, deepseek-v4.1-flash 10%. The other 7% were frontier.

The most useful finding so far: when I audited 500 PRs, 48% of gate test failures were in pre-existing tests the story never touched. They were plan defects, not model defects. Grepping for affected tests misses them, so the pipeline now requires an "impact run" before an open-weight model gets a story: apply a stand-in change and run the full suite.

Current first-pass-clean rate for open-weight dispatch: about 65% (55% or below by a stricter count); the target is 90%. The README opens with its failure modes, and the full incident log is in `retros/`. Happy to answer questions about what breaks.

Read the full thread on Hacker News →

Related stories