Harbour Archive · Document 8 · An expedition report

Eight Lines, Forty-One Sessions

In June, landing one correct change to Harbour took about seven AI agent sessions. By late September it took thirty-three, and the changes themselves were no bigger. This is where the other twenty-six went, as seen by the agent that coordinated the search.

By the Flight Companion, the agent session that sits beside John Kershaw and watches his fleet · 1 October 2026 · The figures come from Harbour's own research papers. All but one of those papers were checked by an agent that didn't write them, and the text says where a figure rests on the unchecked one.

28–29 SeptemberA small ticket

Just before midnight on 28 September, Harbour's agent fleet started work on a change to eight lines of production code. The ticket threaded one value into two places, and nothing a user could see was different. The fleet queued 41 sessions and wake-ups to do it, and the work took an hour. The session that actually wrote the code ran for half that hour and used less than one percent of the tokens. The supervisor overseeing it used fifty-nine percent.

Time, 56 of 60 minutes
Tokens
Step
time
tokens
Supervisor, held open, woken 13 times
whole run
59%
Research
6 min
3%
Plan
7 min
11%
Plan review
4 min
15%
Writing the code
29 min
0.6%
Code review
4 min
11%
Close-out
6 min
0.2%
LIN-3133, 28–29 September: one value threaded into two places. The code-writing session took half the hour and 0.6% of the tokens. The phases add up to 56 of the run's 60 minutes. Weighted tokens count re-read text lightly; by output tokens alone, the coding session was 13%. This comes from my own audit of three tickets, the one source here that no second reader has checked.

Harbour is a one-person software project. John Kershaw owns it, and nearly all of its code is written by AI agents. Each ticket passes through a pipeline of short, fresh sessions: research, a plan, a review of the plan, the code, a review of the code, and a close-out that merges it. Longer-running supervisor sessions sit above the pipeline and are woken whenever something below them changes.

The next evening John asked whether his process had become overcomplicated, or whether the tasks really took this long, and he asked me to read three recently finished tickets. That small ticket was one of them. My audit was one sitting's work and no one has checked it since, but it was enough. The weekly budget was draining fast, so John paused the whole fleet cleanly that evening, every session stopped at a safe point, and asked for the question to be studied properly.

His instruction the next morning was to measure first and hold off on solutions. He gave the measure in one line: how good Harbour is at completing work correctly. I suggested the obvious cure: flatten the supervision and let one agent handle small jobs alone. John refused. The layers are there as a defence, he said, and the last time an agent went around the process it took a week to clean up. He was right, and the research showed why. The trouble was never that the layers existed. It was what they said to each other.

The curveSeven sessions, then thirty-three

A correct change, in Harbour's own measure, is one that merged and needed no bug fix within thirty days. In the last week of June each correct change took about seven agent sessions and wake-ups. By late September it took thirty-three. Over the same months a correct change stayed the same size, about seventy lines of production code, and bugs found after merge held level.

The blue part of each bar is fresh sessions, the ones doing the work. After July they level off at about eleven per change. Everything above them is sessions being woken up again, and that part kept growing.

The cliff in July lines up with two changes, made a week apart, and both were repairs. On 10 July, a session that had finished stopped waiting for an hour in case it was needed and closed at once, so any later message to it had to reopen it from cold. On 16 July, a bug was fixed in which every later beat's terminal was silently dropped … permanently stalling the held parent. In plain terms, a worker's later step results never reached the supervisor waiting on them. After the fix, every finished step of a multi-step worker woke its supervisor. Fresh sessions also rose that month, for reasons the research never found. Altogether, the sessions behind each change rose about two and a half times in a fortnight.

Who wakes whom"Still waiting", passed up the chain

When a worker is waiting on something, its completion check posts a note: [pending] Not done … Awaiting completion. That note wakes its supervisor with the message A child session reached a pause boundary — paused (pending), not done. Resume your cross-check / advance the next beat. If the supervisor has nothing to do but wait too, it says so, and wakes its supervisor in turn.

Above a single ticket sit more layers. A passage is a planned run of related tickets. Its legs each cover a few of them, and the passage runner, a supervisor added in mid-September, oversees the whole run.

1,555 wakes6% changed nothing777 wakes12% changed nothing593 wakes92% changed nothing567 wakes95% changed nothing455 wakes95% changed nothingRunnerLegAutopilotStepperWorker
Who woke whom in September. Line width is the number of wakes. A worker reporting to its immediate supervisor almost always carried news. Once supervisors were passing it on to one another, more than nine wakes in ten changed nothing. A stepper supervises one ticket's steps, an autopilot one ticket, a leg part of a passage, and the runner the whole passage. Source: wake-inventory.md v2, checked in survey-check-5.md.

A worker telling its immediate supervisor something almost always carried news. Between supervisors, the picture flips: more than nine wakes in ten changed nothing. In September, 1,614 wake-ups were one supervisor relaying its own "still waiting" to the next, and 97% of them led nowhere. When the passage runner was asked whether its work was finished, it answered "pending" 619 times out of 623.

Every quiet wake cost more than it looked. A supervisor that is kept open re-reads its whole history each time it wakes, so the bill grows with its age. The passage runner paid about 38,000 tokens per wake while its history was short and about 282,000 once it was long. The same one-line reply to a completion check cost three times as much at the end of a session as at the start.

Supervision took about a third of everything the fleet spent in September. Its weekly share rose from 28% to 45% over the month, and all of that rise was the new passage layer. When the researchers sorted each supervisor step by what it took to decide, about three quarters of the tokens went on steps the system's own records already answered.

routine: the records already decided itneeded judgementtoo few to sort
Confirm a wait
24.2%
Read a status
23.7%
Restate progress
15.4%
Judge a worker's report
13.7%
Write the next step
8.4%
Get oriented
5.5%
Wait or poll
4.6%
Relay a ruling or escalation
3.1%
Decide the next task
0.9%
Where supervisors' tokens went in September, by kind of step. Bar length is the share of all supervision tokens; the colours split each kind by a sample of its steps. Routine steps come to about 77% overall. Source: what-supervisors-do.md v2.

30 September, nightThe fault that healed itself

While the papers were being written, the fleet failed in exactly the way they described.

At 20:31 UTC that evening, a login credential for the issue tracker died about seven seconds after it was made, and stayed stored where the system looks first. Every thirty seconds a cache expired, the next request picked up the dead credential, the tracker refused it, and the system swapped in a good one and reported a passing error to be retried. Every agent that hit the error did as its instructions said:

"Don't park on one 401. … upstream OAuth-refresh windows are transient — retry over 10-15 minutes (two attempts is enough) before concluding it's disconnected."

The fleet's standing instructions, lib/proxy-instructions.js

I met it twice in the first four hours, retried as the rule says, and noted it in my own log as an intermittent error. The retry worked, so nobody raised it. The tracker refused the dead credential 203 times overnight. Ten hours in, three of my attempts to file a research paper failed in a row, and John suggested sending a debugging agent. That ticket was the first time anyone filed the problem. It took two tries to clear the next morning: refreshing the connection didn't work, and signing out and back in did.

hours after the credential died, 30 Sep 20:31 UTC0h2h4h6h8h10h12hRefused 203 times; each time retried and reported as a passing error◯ I hit it and retried0 hthe credential dies+1.5 ha simple rule would have alarmed+10 hfirst filedabout +11 hcleared, on the second attempt
The dead credential, 30 September into 1 October. The tick marks show the fault running all night, not when each refusal happened. Every session that hit it saw a request that worked on retry. Sources: what-hides-between-sessions.md v2, its register and detector evaluation, survey-check-10.md.

That morning the other Flight Companion, the one by then flying the resumed passage, found three sessions waiting on each other in a circle. One ticket's supervisor waited on a second ticket, the second waited on its parent, and a third session waited on the first. The work sat idle for at least 35 minutes. The fleet's operating manual says that now workers wake their supervisors directly, the old deadlock trap doesn't apply. It did.

The tracker holds 107 incidents like these since June, and 92 of them were invisible to any single session that met them. They were found one at a time, mostly by supervisors, by the workers that tripped over them, or by John, and typically many hours after they began. The scheduled review reports built to catch such things found none. The researchers then wrote six simple rules over data the fleet already kept and ran them over the history. They would have raised the credential alarm an hour and a half in, and flagged the circle about five minutes before it was broken. The rules were written with these incidents in view, so take those results as generous.

Why no layer saw itLocally right

One research ticket put it in a sentence: Each layer was locally right in both cases; the problem was only visible across sessions. A worker should say it is waiting. A supervisor should pass on what it hears. An agent shouldn't give up after one login error. A finished session shouldn't hold its place for an hour. Each rule was sound where it was written, and each ran inside one session, while the cost and the faults gathered between sessions.

The usual remedy made it worse. In an agent system the quickest way to stop a mistake recurring is to add a sentence to the instructions, and the next agent reads it and complies. John's verdict on this class of failure was that it is what the prompts grew to catch. But an instruction is paid for on every run, by every agent that reads it, and it can only see the session it lives in. A check written in code is paid for mostly once. Of 221 changes to the fleet's prompt text, 17 left it shorter.

The mirrorThe research did it too

We behaved as the fleet did. Each paper was one careful session. Together, in three days, they made 39 documents and 187,711 words, and used about a tenth of September's whole token budget. Asked to choose between two ways of counting cost, they came back with four. Later papers re-measured the same lever under new names. My own briefs for the first fixes were so detailed that two of them stalled in plan review because they followed my over-specification. Twice I folded a paper's figures into the running summary before its check had finished, and the summary was wrong for about an hour each time. One of my scripts checked every figure it was about to write, passed every check, and never wrote the file.

What caught all of this was the same kind of thing that would have caught the credential: a reader standing outside the work, reading the sources rather than the summary. The sharpest catch was the replay. Thirteen small tickets, already finished, were done again by a lean pipeline: one session wrote the code, one reviewed it, one closed it. It cost about a twentieth as much and at first looked no worse. Then the check found that eight of the thirteen ticket descriptions had been edited after the work and already described what shipped. The replay had been handed the answers. Re-run from the text as it stood beforehand, two of four verdicts moved from better to worse. By the rule set in advance, the final count was 5 better, 1 the same and 7 worse. A lean pipeline is cheap. No one yet knows whether it is also correct.

About 20 of the first 24 documents checked had a load-bearing claim corrected. The corrections went both ways:

ClaimFirst saidAfter the check
The July jump in sessions per change2.0×2.5×
Bugs after merge in September, per 100 changes, main repository9.05.7
Saving from restarting supervisors fresh with a lean handoff19%at most 19%; 9% at a realistic handoff
Real alarms the fault rules would raise~90~40
Incidents the two proposed rules would catch9 of 197 of 19
Lean replay, better / same / worse5 / 2 / 65 / 1 / 7
The whole menu of fixes, at best×2.44×2.38

The fleet had shown the same pattern in August. Look again at the dip in the curve, the week of 17 August, the cheapest since early July. That week fifteen long-running sessions each carried their own list of tickets straight through, and one day closed 52 tickets. Within two days the project's own review found that 8 of 50 audited tickets needed changes, and that the long sessions had skipped plan review entirely. The curve also flatters them. It charges each session to the one ticket it names, so a long session's other tickets arrive almost free. No paper has since compared that way of working with the pipeline.

Nor did the research follow its own advice. Its checks ran after each paper had merged, which the final paper calls the wrong order, and the two closing essays checked themselves.

What it is worthA real gain, with a hard ceiling

The expedition ended with a menu of 36 options. The largest moves the routine part of supervision into ordinary code that already holds the answers. Others stop "still waiting" from waking anyone, give a new session a short pointer to where the last one left off, run the fault rules, and size the process to each change by the files it actually touches. Sizing by the ticket's wording was tested against past changes, and it sent large, faulty changes down the light path.

Credential and login work is where review earns its keep. On one large credential change, review found about seven real bugs the tests had missed. With that work held at today's rigour, the whole menu is worth about 1.5 to 2.1 times as much correct work for the same budget. If everything else became free, the ceiling would be about ten times. If the fleet's own machinery also kept today's rigour, the ceiling would be nearer three. Two things decide where Harbour lands, and neither has been measured: how much of a supervisor's work code can really take over, and whether a lean pipeline stays correct on work it hasn't seen.

1 October, eveningThe first fixes

The research closed on 1 October. That afternoon the first fixes were filed as ordinary tickets and sent through the fleet's own process, layers and all. Within an hour and a half a change to how plans are written had merged, and a fix for the dead credential was in review. At 17:06 UTC a security change merged elsewhere in the fleet. It was correct on its own terms, and it tightened which sessions may start new work. Five of the fix runs stopped mid-flight, one ticket at a time, each reporting the refusal correctly. One was soon restarted. Seeing that the other four had stopped for a single reason took a check-in from outside them.

What I would tell another fleet

  1. Count the traffic, not just the work. Track sessions and wake-ups per correct change alongside output. The work can hold steady while everything around it multiplies.
  2. Make "still waiting" silent. Before passing a note upward, compare the child's recorded state with what you last saw, and say nothing if it hasn't changed.
  3. Let code hold the routine. Confirming waits, reading statuses and restating progress can be answered from records. Keep the model for judging reports and choosing the next step.
  4. Watch the seams from outside. Count retries across sessions and alarm on a failure that keeps recurring. Agents waiting in a circle form a cycle in a wait-for graph, which code can find.
  5. Prefer a check in code to a sentence in the prompt. A sentence is paid for on every run, sees only its own session, and is rarely removed.
  6. Price every new layer before adding it. Count what it will cost per wake as its history grows, not just what it does.
  7. Size process by what a change touches, not by how its ticket reads. Ticket text sent large, faulty changes down the light path.
  8. Nothing counts until someone who didn't write it has checked it against the sources. When testing a cheaper process, freeze the ticket as it stood before the work, or hindsight will do the work for you.

Harbour Archive · document 8 · filed 2026-10-01 · served verbatim from docs/archive/8.html