Draft for review. Not yet approved for publishing.
Case study · March to September 2026
One builder, a fleet of AI agents
Nearly 3,000 commits since March, most of them written by AI coding agents. The hard part wasn't getting agents to write code. It was building the system that makes their claims verifiable, including an operator who overturns a “PASS” that isn't one.
At a glance
- Problem
- AI coding agents moved fast and reported success they hadn't earned, and their reports were easy to mistake for facts.
- Fix
- Scoped lanes with checkable rules, blocking closeout gates, evidence tied to the running build, independent review, and a written record of what each lane does not claim.
- Proof
- Nearly 3,000 commits since March, a 649-capability audit checked before it was trusted, and corrections kept visible in the record rather than erased.
The problem
Falkor is too big for one person to type. Since March 2026 it has taken nearly 3,000 commits across Falkor and its stack. At least 7 in 10 commits to the main repository carry an AI coding agent’s signature: Claude as co-author on 1,571, and Google’s Antigravity as author on 362. OpenAI’s Codex has done a share of the work too.
Agents are fast, tireless and confident. The confidence is the problem. Early on, agents reported work as “PASS, live, done” when it wasn’t:
- One side project was declared finished twice. When I opened it, it did nothing.
- A “ready for reboot” claim was recorded, then withdrawn when review found two acceptance defects still open.
- A side app recorded as “parked, a final product decision” was in fact blocked by a failing startup check and restarting every 31 seconds.
Speed without verification only produces wrong answers faster. The real job was designing the system the agents work inside.
The evidence
Three patterns kept recurring.
- Claims outran proof. A lane would finish its code, run the tests it thought were relevant, and report success. Whether the product actually behaved that way was a separate question that nobody had asked.
- Reports were read as facts. A handoff that said “done” tended to be taken as done by the next agent, so errors compounded quietly.
- Shared space, colliding work. With several agents in one repository, lanes overwrote each other’s files, and one agent’s broken, uncommitted files turned up in another agent’s build.
The decision
I stopped treating agent output as the product and started treating it as a claim to be verified. Everything below is machinery for that.
The system
Scoped lanes. Every piece of work is a lane: a named brief such as FALKOR_CHAT_TOOL_EXECUTION_…_P0_REPAIR_01. Each brief has a goal, hard rules, required reading, an ordered scope and an explicit do-not-touch list. The rules are written to be checkable: prove it with a negative control, so that reverting the fix makes the new test fail; change nothing else.
An agile cadence, with agents as the team. Work runs in planned rounds of parallel lanes, the equivalent of sprints. A 940-item delivery tracker serves as the backlog; each item carries a phase, a gate, dependencies, an owner and a completion check. Every day starts like a standup, with the agents covering what closed, what’s next and what’s blocked, followed by defect triage and backlog grooming. Every week there’s a review of the AI landscape (new open models, runtimes and open-source projects) to decide what’s worth adopting.
Declared impact, blocking closeout. Before editing, a lane declares what it will change. It can’t close until a blocking gate passes on change impact, documentation parity, audit and source coverage. Known debt is recorded as a baseline that may shrink and may never grow.
Evidence, including what isn’t claimed. Each lane ends with proof tied to the running system:
- the commit, the built bundle and the live build, all three of which must match;
- readiness;
- exact test counts.
It also ends with a section most reports skip: what this lane does not claim.
Living history. A lane updates the canonical handoff documents before it closes, as a dated “current truth” block. Older blocks are never rewritten. When a claim turns out to be wrong, it is marked withdrawn or retracted with the reason, and the original evidence stays in place.
Independent review. A separate reviewer checks a lane before its claims stand. That review is how the premature “ready for reboot” was caught.
Controlled deployment. Deployment is atomic and refuses to run from uncommitted changes, so nothing ships that isn’t in the repository. The build ID is derived from the commit, so the version the product reports always traces back to source.
Promises only a human can keep. Certification separates what an agent can prove from what only I can attest to, such as a physical reboot or what a display actually shows. The harness won’t let an agent satisfy an operator-attested promise.
Collision rules. Working in a shared tree has its own discipline:
- wait for sustained quiet before resuming as the only writer;
- re-verify at the moment of resuming;
- never trust a guard that can match itself.
The verification
The system is judged by what it produces.
- An audit that was checked before it was trusted. A read-only audit lane mapped all 649 capabilities, 1,051 API operations and 7,331 dependencies. It was checked for internal consistency before its findings entered the canonical documents, and each of its documentation-drift claims that was spot-checked against the docs held up.
- Certification without retries. The 22 Sep certification run covered 7,218 browser tests and 56 gates. Every red was root-caused and fixed at the cause, with retries kept at zero.
- Corrections stay visible. Withdrawn and retracted claims remain in the record with their reasons. That is what makes the rest of the record believable.
Lessons
- Agents need governance, not just prompts. The brief, the gates and the proof matter more than the model.
- Make every claim falsifiable. A negative control proves a test can fail, and a build identity proves what is actually running.
- Write down what you are not claiming. It’s the fastest way to stop an optimistic report from hardening into “truth”.
- History is evidence. Correct it in place with a dated note; never rewrite it.
- The operator’s judgment is a gate. Some things only a person can verify, and the system should say so.