AICodingAgentHarness: a governance layer for coding agents

a computer screen with a bunch of code on it

Coding agents write code. They don’t decide whether that code deserves to ship. That gap is the problem AICodingAgentHarness is built to close.

The open-source repo, published by a Lead AI Architect at Microsoft, wraps coding agents in a structured pipeline: named agents per stage, human approval gates before consequential steps, and a dedicated evaluator agent that is explicitly separate from the builder. The harness lives in the repo itself as real files agents read automatically, not in a prompt library.

‍ How the pipeline works

Each stage has one agent, one output artifact, and one clear responsibility:

  • problem-statement-creation: separates confirmed requirements from assumptions and open questions. Output goes to output/PROBLEMSTATEMENT.md. A human reviews before architecture starts.
  • technical-architect: runs the draft design through five internal personas (Architect, Critic, Azure Pragmatist, Security Reviewer, Operability/Cost Reviewer) and writes unresolved decisions to output/TechnicalGaps.md rather than silently assuming them. Three checkpoints gate this stage before anything downstream starts.
  • implementation-planner: every task in the output plan answers five questions: who owns it, which files they may touch, which contracts it must follow, which acceptance criteria it satisfies, and what evidence proves it’s done.
  • parallel-build-orchestrator: fans out to frontend, backend, AI, and data engineer subagents when file ownership doesn’t overlap. Stages only serialize when there’s a real dependency. A one-file bug fix skips all of this and routes to a single coding-agent instead.
  • code-reviewer + verification-evaluator: the evaluator is structurally separate from the builder. It checks results against the approved problem statement, design, plan, and test evidence, then returns ITERATE or PASS. Deleting or weakening a test to force a PASS is a blocking violation. The iteration loop is bounded with a budget and human escalation when progress stalls.
  • agent-feedback: triggers on PASS only, captures reusable lessons with evidence and scope, and holds them for human review before promoting to memory. Future agents only receive memories relevant to their task and path.
lines of HTML codes

What the controlled study showed

The author evaluated the harness on a fixed ConverseHub build across multiple configurations. These are not industry benchmarks; they’re controlled implementation results tracked in benchmarks/:

  • Quality score: 6.64 to 9.02 (targeted repository memory vs. bare run)
  • Output tokens: 62.3K to 49.0K (continuous full harness vs. bare run)
  • Tool calls: 83 to 52
  • Isolated-stage sessions: 170.9K tokens, roughly 2.7x the continuous run

That last number is the one worth flagging. When every stage started in a fresh session, output ballooned. Role separation helped quality, but context amnesia was expensive. The author’s conclusion: a harness has to earn its orchestration cost through better quality and less rediscovery, not assumed to save tokens by default.

Four harness tiers, not four levels of ceremony

The repo defines four concrete tiers matched to risk and coordination need. A one-file bug fix routed through the full pipeline pays for parallel lanes and an evaluator loop it doesn’t need. A cross-module feature routed through coding-agent alone skips the design checkpoint that would have caught a bad assumption before code existed. The tier is a judgment call made once, up front, based on how many files and owners the change touches.

Start with AGENTS.md (about five minutes to read), then .github/agents/technical-architect.agent.md for a fully worked spec, then gan-harness/ to see a generator-evaluator loop on disk.

Stay on top of AI & Automation with BizStack Newsletter
BizStack  —  Entrepreneur’s Business Stack
Logo