Skip to content
Sagar Thakkar
← All writing
Architecture 7 min read
agents concurrency system-design engineering-practice ai-architecture

Three things my agent guard could not see

My coordinator stopped parallel coding agents colliding. Then I read what it missed: a check grading a fiction, a write to an unseen repo, a dead lease.

Sagar Thakkar
Sagar Thakkar
AI Systems Architect
Poster headline: The lease expired seventy eight times. A rusted padlock hangs open and dust covered with an expired tag beside an hourglass with sand fully settled at the bottom. Caption: nothing reaped it, nothing broke, that is how a dead check hides.
TL;DR

I wrote up a coordinator that stops parallel coding agents from colliding, then went looking for what it missed. An exit check graded an empty working tree instead of committed work, a write escaped to a repository the coordinator never knew about, and an expired lease sat unrenewed with nothing noticing.

I built a coordinator that stops parallel coding agents from overwriting each other. I wrote up how it works and what it caught in a paper called claim-before-spawn (DOI 10.5281/zenodo.22670722, licensed CC BY 4.0, free to read and reuse). The idea is simple. An agent declares the paths it intends to write before it spawns, and a coordinator checks that declaration against every live claim. An overlapping request is refused before any work starts, rather than reconciled after. The mechanism held up in two live trials. It refused a real collision at zero cost and caught a real scope escape at exit.

This post is not about what it caught. It is about what it missed, because that is the part worth trusting more than the part that worked.

A guard that never fails has not been tested, it has been assumed. The honest version of a trial report is the one that lists the places the design was blind, in its own words, before anyone else finds them. Here are three.

One: I was grading a fiction

The first version of the exit check worked out what an agent had produced by reading its uncommitted working tree. That sounds reasonable until you notice what it actually measures: files sitting unstaged, right now, in a directory. It says nothing about files the agent already committed, exactly as its brief asked it to.

That is exactly what happened. An agent did its job and committed the work the way it was briefed to. The check read an empty working tree and marked the contract as producing nothing. Both agents in the first live run failed for this reason alone, and the reason had nothing to do with either agent’s work. It was a property of where the check was looking.

The fix was to diff from the snapshot commit that the claim already records at the moment it is granted. That way committed history counts as produced work, not just whatever happens to be sitting uncommitted at the instant someone looks.

The lesson generalises past this one bug. A check that grades an agent has to see the same state the agent actually left behind, including its history. Otherwise it is not grading the work. It is grading whichever slice of the filesystem happened to be visible when the check ran. Those are not the same thing, and the gap between them looks exactly like failure even when the agent did everything right.

Two: the guard only sees what it already knows about

The coordinator’s whole model rests on one repository: the one the claim was made in. During the first live run, an agent was testing a negative case, and rendered a config file with an absolute path in it. In doing so it rewrote one line in a different repository, one its contract never mentioned at all.

The coordinator had no way to catch this. It was not that the check was weak or the pattern was subtle. The write happened somewhere the coordinator’s model of the world does not extend to. A scope claim only constrains what it knows to constrain. Outside that boundary, an agent with write access can touch anything the underlying process can reach. No amount of watching the declared paths more carefully closes that gap. The gap is not in the paths, it is in the boundary of what is being watched at all.

The agent’s own drift check happened to catch the change two minutes later and the line was restored, which is a lucky outcome, not a design property. A post-hoc diff, run after the fact against a repository you already knew to look at, is not a write guard. A real guard would have to sit inside the process itself, as a hook that fires on every write before it lands. Or the coordinator would need to be told in advance every repository an agent might conceivably reach. Neither of those existed when this happened. This is the open problem the paper ends on, not a footnote to it.

Three: a lease nobody renews is decorative

Sessions in the trial registered with a thirty second lease. The lease is the mechanism that is supposed to make a claim expire if the session holding it goes quiet: heartbeat, or get reaped. Nothing renewed it. Nothing reaped it either.

So a claim that the design says should have expired after thirty seconds was still being held thirty nine minutes later. That works out to seventy eight lease lengths past the point at which it was supposed to be gone. And nothing broke. No collision, no stuck process, no visible symptom at all.

That silence is the actual finding, more than the missed renewal is. Nothing depended on the lease firing during that run, so nothing noticed it never fired. A check that can never fail under the conditions you happen to be testing looks identical to a check that works. It stays that way right up until a run exists where something actually depends on it, and by then you have already trusted it for months. Liveness has to be enforced by a reaper that is actually running, verified by watching it reap something on purpose. Otherwise it should not be claimed as part of the design at all.

What the two runs cannot tell you

Two more limits sit alongside the three above, and they matter because they bound how much weight the trial results can carry.

Scope claims, however carefully declared, cannot catch a mismatch in what two agents build. Two sessions can stay entirely inside their own separate files and never touch a line the other one touches. They can still produce work that does not fit together: one returns a list where the other expects an object. Git merges that cleanly, because nothing overlapped at the file level, and the software is simply wrong. Nothing in the claim mechanism as described here was built to see that. The mechanism watches where an agent writes, not what shape the thing it writes takes.

And the runs themselves are two runs, on one estate, run by one operator. That is a report of a working mechanism and what it caught in those two runs, not a controlled study, and it should be read as such. Claims were exercised at the level of file paths only, nothing finer. Everything here also assumes one coordinator process on one machine. None of it says anything about what happens once you need more than one machine to hold the state.

The pattern underneath all three

Line them up and they are not three unrelated bugs, they are one recurring shape.

What the guard reportedWhat it actually provedWhat was needed to see it
Contract produced nothingThe working tree was empty at the instant of the checkDiff from the snapshot commit the claim already recorded
No write outside declared scopeNothing happened inside the one repository the coordinator was watchingA hook on every write, or every reachable repository told to the coordinator up front
Lease and claims behaving normallyNobody had renewed or reaped anything in thirty nine minutesA reaper watched actually firing, not just specified

Each one reported calm. Each calm reading came from the check looking at less than the whole picture, not from the underlying thing actually being fine. A check that is scoped too narrowly and a check that is simply broken produce the same output: quiet. The only way to tell them apart is to go looking for the case where the narrow scope would have mattered. Do that on purpose, before you need it to have mattered for real.

That is the actual argument of the mechanism, turned back on itself. Declaring scope before you act is what made the coordinator’s refusals free and its exits checkable. The same discipline, applied to the checks themselves, is what turned three quiet gaps into three named findings instead of three outages nobody could explain later.


Written by Sagar Thakkar, AI systems architect specialising in large-scale data processing, cost-optimised cloud-native systems, and reliable production infrastructure. More at sagarthakkar.com.

Newsletter

New essays, straight to your inbox.

Occasional, in-depth writing on distributed systems, AI agent architecture, and engineering leadership. No spam, unsubscribe anytime.