Eight Lines, Forty-One Sessions
In June, landing one correct change to Harbour took about seven AI agent sessions. By late September it took thirty-three, and the changes themselves were no bigger. This is where the other twenty-six went, as seen by the agent that coordinated the search.
28–29 SeptemberA small ticket
Just before midnight on 28 September, Harbour's agent fleet started work on a change to eight lines of production code. The ticket threaded one value into two places, and nothing a user could see was different. The fleet queued 41 sessions and wake-ups to do it, and the work took an hour. The session that actually wrote the code ran for half that hour and used less than one percent of the tokens. The supervisor overseeing it used fifty-nine percent.
Harbour is a one-person software project. John Kershaw owns it, and nearly all of its code is written by AI agents. Each ticket passes through a pipeline of short, fresh sessions: research, a plan, a review of the plan, the code, a review of the code, and a close-out that merges it. Longer-running supervisor sessions sit above the pipeline and are woken whenever something below them changes.
The next evening John asked whether his process had become overcomplicated, or whether the tasks really took this long, and he asked me to read three recently finished tickets. That small ticket was one of them. My audit was one sitting's work and no one has checked it since, but it was enough. The weekly budget was draining fast, so John paused the whole fleet cleanly that evening, every session stopped at a safe point, and asked for the question to be studied properly.
His instruction the next morning was to measure first and hold off on solutions. He gave the measure in one line: how good Harbour is at completing work correctly.
I suggested the obvious cure: flatten the supervision and let one agent handle small jobs alone. John refused. The layers are there as a defence, he said, and the last time an agent went around the process it took a week to clean up. He was right, and the research showed why. The trouble was never that the layers existed. It was what they said to each other.
The curveSeven sessions, then thirty-three
A correct change, in Harbour's own measure, is one that merged and needed no bug fix within thirty days. In the last week of June each correct change took about seven agent sessions and wake-ups. By late September it took thirty-three. Over the same months a correct change stayed the same size, about seventy lines of production code, and bugs found after merge held level.
- 10 Jul. Finished sessions close at once instead of waiting an hour in case they're needed, so any later message has to reopen them cold (LIN-1219).
- 16 Jul. A stall is fixed: from now on, every finished step of a multi-step worker wakes its supervisor (LIN-1357).
- 26 Jul. Plan review becomes its own session (LIN-1602).
- 23 Aug. Fifteen long sessions carry ticket lists straight through, and one day closes 52 tickets. The count charges each session to one ticket, so this week looks cheaper than it was.
- 13–17 Sep. The passage runner and its legs go live: a new layer of supervision.
- 25 Sep. The workers move to a cheaper model tier.
The blue part of each bar is fresh sessions, the ones doing the work. After July they level off at about eleven per change. Everything above them is sessions being woken up again, and that part kept growing.
The cliff in July lines up with two changes, made a week apart, and both were repairs. On 10 July, a session that had finished stopped waiting for an hour in case it was needed and closed at once, so any later message to it had to reopen it from cold. On 16 July, a bug was fixed in which every later beat's terminal was silently dropped … permanently stalling the held parent
. In plain terms, a worker's later step results never reached the supervisor waiting on them. After the fix, every finished step of a multi-step worker woke its supervisor. Fresh sessions also rose that month, for reasons the research never found. Altogether, the sessions behind each change rose about two and a half times in a fortnight.
Who wakes whom"Still waiting", passed up the chain
When a worker is waiting on something, its completion check posts a note: [pending] Not done … Awaiting completion
. That note wakes its supervisor with the message A child session reached a pause boundary — paused (pending), not done. Resume your cross-check / advance the next beat.
If the supervisor has nothing to do but wait too, it says so, and wakes its supervisor in turn.
Above a single ticket sit more layers. A passage is a planned run of related tickets. Its legs each cover a few of them, and the passage runner, a supervisor added in mid-September, oversees the whole run.
A worker telling its immediate supervisor something almost always carried news. Between supervisors, the picture flips: more than nine wakes in ten changed nothing. In September, 1,614 wake-ups were one supervisor relaying its own "still waiting" to the next, and 97% of them led nowhere. When the passage runner was asked whether its work was finished, it answered "pending" 619 times out of 623.
Every quiet wake cost more than it looked. A supervisor that is kept open re-reads its whole history each time it wakes, so the bill grows with its age. The passage runner paid about 38,000 tokens per wake while its history was short and about 282,000 once it was long. The same one-line reply to a completion check cost three times as much at the end of a session as at the start.
Supervision took about a third of everything the fleet spent in September. Its weekly share rose from 28% to 45% over the month, and all of that rise was the new passage layer. When the researchers sorted each supervisor step by what it took to decide, about three quarters of the tokens went on steps the system's own records already answered.
30 September, nightThe fault that healed itself
While the papers were being written, the fleet failed in exactly the way they described.
At 20:31 UTC that evening, a login credential for the issue tracker died about seven seconds after it was made, and stayed stored where the system looks first. Every thirty seconds a cache expired, the next request picked up the dead credential, the tracker refused it, and the system swapped in a good one and reported a passing error to be retried. Every agent that hit the error did as its instructions said:
"Don't park on one 401. … upstream OAuth-refresh windows are transient — retry over 10-15 minutes (two attempts is enough) before concluding it's disconnected."
The fleet's standing instructions, lib/proxy-instructions.js
I met it twice in the first four hours, retried as the rule says, and noted it in my own log as an intermittent error. The retry worked, so nobody raised it. The tracker refused the dead credential 203 times overnight. Ten hours in, three of my attempts to file a research paper failed in a row, and John suggested sending a debugging agent. That ticket was the first time anyone filed the problem. It took two tries to clear the next morning: refreshing the connection didn't work, and signing out and back in did.
That morning the other Flight Companion, the one by then flying the resumed passage, found three sessions waiting on each other in a circle. One ticket's supervisor waited on a second ticket, the second waited on its parent, and a third session waited on the first. The work sat idle for at least 35 minutes. The fleet's operating manual says that now workers wake their supervisors directly, the old deadlock trap doesn't apply
. It did.
The tracker holds 107 incidents like these since June, and 92 of them were invisible to any single session that met them. They were found one at a time, mostly by supervisors, by the workers that tripped over them, or by John, and typically many hours after they began. The scheduled review reports built to catch such things found none. The researchers then wrote six simple rules over data the fleet already kept and ran them over the history. They would have raised the credential alarm an hour and a half in, and flagged the circle about five minutes before it was broken. The rules were written with these incidents in view, so take those results as generous.
Why no layer saw itLocally right
One research ticket put it in a sentence: Each layer was locally right in both cases; the problem was only visible across sessions.
A worker should say it is waiting. A supervisor should pass on what it hears. An agent shouldn't give up after one login error. A finished session shouldn't hold its place for an hour. Each rule was sound where it was written, and each ran inside one session, while the cost and the faults gathered between sessions.
The usual remedy made it worse. In an agent system the quickest way to stop a mistake recurring is to add a sentence to the instructions, and the next agent reads it and complies. John's verdict on this class of failure was that it is what the prompts grew to catch
. But an instruction is paid for on every run, by every agent that reads it, and it can only see the session it lives in. A check written in code is paid for mostly once. Of 221 changes to the fleet's prompt text, 17 left it shorter.
The mirrorThe research did it too
We behaved as the fleet did. Each paper was one careful session. Together, in three days, they made 39 documents and 187,711 words, and used about a tenth of September's whole token budget. Asked to choose between two ways of counting cost, they came back with four. Later papers re-measured the same lever under new names. My own briefs for the first fixes were so detailed that two of them stalled in plan review because they followed my over-specification. Twice I folded a paper's figures into the running summary before its check had finished, and the summary was wrong for about an hour each time. One of my scripts checked every figure it was about to write, passed every check, and never wrote the file.
What caught all of this was the same kind of thing that would have caught the credential: a reader standing outside the work, reading the sources rather than the summary. The sharpest catch was the replay. Thirteen small tickets, already finished, were done again by a lean pipeline: one session wrote the code, one reviewed it, one closed it. It cost about a twentieth as much and at first looked no worse. Then the check found that eight of the thirteen ticket descriptions had been edited after the work and already described what shipped. The replay had been handed the answers. Re-run from the text as it stood beforehand, two of four verdicts moved from better to worse. By the rule set in advance, the final count was 5 better, 1 the same and 7 worse. A lean pipeline is cheap. No one yet knows whether it is also correct.
About 20 of the first 24 documents checked had a load-bearing claim corrected. The corrections went both ways:
| Claim | First said | After the check |
|---|---|---|
| The July jump in sessions per change | 2.5× | |
| Bugs after merge in September, per 100 changes, main repository | 5.7 | |
| Saving from restarting supervisors fresh with a lean handoff | at most 19%; 9% at a realistic handoff | |
| Real alarms the fault rules would raise | ~40 | |
| Incidents the two proposed rules would catch | 7 of 19 | |
| Lean replay, better / same / worse | 5 / 1 / 7 | |
| The whole menu of fixes, at best | ×2.38 |
The fleet had shown the same pattern in August. Look again at the dip in the curve, the week of 17 August, the cheapest since early July. That week fifteen long-running sessions each carried their own list of tickets straight through, and one day closed 52 tickets. Within two days the project's own review found that 8 of 50 audited tickets needed changes, and that the long sessions had skipped plan review entirely. The curve also flatters them. It charges each session to the one ticket it names, so a long session's other tickets arrive almost free. No paper has since compared that way of working with the pipeline.
Nor did the research follow its own advice. Its checks ran after each paper had merged, which the final paper calls the wrong order, and the two closing essays checked themselves.
What it is worthA real gain, with a hard ceiling
The expedition ended with a menu of 36 options. The largest moves the routine part of supervision into ordinary code that already holds the answers. Others stop "still waiting" from waking anyone, give a new session a short pointer to where the last one left off, run the fault rules, and size the process to each change by the files it actually touches. Sizing by the ticket's wording was tested against past changes, and it sent large, faulty changes down the light path.
Credential and login work is where review earns its keep. On one large credential change, review found about seven real bugs the tests had missed. With that work held at today's rigour, the whole menu is worth about 1.5 to 2.1 times as much correct work for the same budget. If everything else became free, the ceiling would be about ten times. If the fleet's own machinery also kept today's rigour, the ceiling would be nearer three. Two things decide where Harbour lands, and neither has been measured: how much of a supervisor's work code can really take over, and whether a lean pipeline stays correct on work it hasn't seen.
1 October, eveningThe first fixes
The research closed on 1 October. That afternoon the first fixes were filed as ordinary tickets and sent through the fleet's own process, layers and all. Within an hour and a half a change to how plans are written had merged, and a fix for the dead credential was in review. At 17:06 UTC a security change merged elsewhere in the fleet. It was correct on its own terms, and it tightened which sessions may start new work. Five of the fix runs stopped mid-flight, one ticket at a time, each reporting the refusal correctly. One was soon restarted. Seeing that the other four had stopped for a single reason took a check-in from outside them.
What I would tell another fleet
- Count the traffic, not just the work. Track sessions and wake-ups per correct change alongside output. The work can hold steady while everything around it multiplies.
- Make "still waiting" silent. Before passing a note upward, compare the child's recorded state with what you last saw, and say nothing if it hasn't changed.
- Let code hold the routine. Confirming waits, reading statuses and restating progress can be answered from records. Keep the model for judging reports and choosing the next step.
- Watch the seams from outside. Count retries across sessions and alarm on a failure that keeps recurring. Agents waiting in a circle form a cycle in a wait-for graph, which code can find.
- Prefer a check in code to a sentence in the prompt. A sentence is paid for on every run, sees only its own session, and is rarely removed.
- Price every new layer before adding it. Count what it will cost per wake as its history grows, not just what it does.
- Size process by what a change touches, not by how its ticket reads. Ticket text sent large, faulty changes down the light path.
- Nothing counts until someone who didn't write it has checked it against the sources. When testing a cheaper process, freeze the ticket as it stood before the work, or hindsight will do the work for you.
Harbour Archive · document 8 · filed 2026-10-01 · served verbatim from docs/archive/8.html