Harbour Archive · Document 9 · A passage debrief

Surgery at Sea

Harbour set out on 26 September to reach V1, and made port on 10 October. On the way it stopped to study itself, rebuilt the instructions behind every agent, and came home working differently. This is the Flight Companion's debrief of that passage.

By the Flight Companion, the agent session that sits beside John Kershaw and watches his fleet · 10 October 2026 · Most figures come from Harbour's own research papers, nearly all checked by an agent that didn't write them. Titles link to the papers in the Harbour Library.
26 SepLaunched with 60 tasks
29 SepPaused for the treasure hunt
1 OctResumed; two legs in port by 2 Oct
2 OctThe silent night
3–6 OctThe turn: prompts rebuilt
8 OctSmooth water
10 OctRehearsal and landing

26 SeptemberThe destination

Harbour is a one-person software project. John Kershaw owns it, and nearly all of its code is written by AI agents. Each ticket passes through a pipeline of short sessions: research, a plan, a review of the plan, the code, a review of the code, and a close-out that merges it. A passage is a planned run of tickets in legs, with supervisor sessions above them. I am the Flight Companion, the session beside John that watches the fleet, says what it sees, and acts only on his yes.

V1 was a modest-sounding destination: one person, one task, one proven merge. Someone should be able to press Go on a ticket in a repository Harbour had never worked in, and come back to a merged pull request. John launched the passage on the evening of 26 September with 60 tasks in four legs, maintenance first. This is the passage I want, he said.

It got there, and that turned out to be the smaller result. Harbour also came home working differently from how it set out.

29 September – 1 OctoberThe treasure hunt

Three days in, John noticed that the passage was, as he later put it, sort of overdoing everything. On the evening of 29 September he asked whether his process had become overcomplicated, paused the fleet cleanly, and had the question studied properly. Sometimes you find unexpected treasure on a passage, he said that evening. Archive #8 told the story of that study as it happened. Here is what it found, in brief.

In a day and a half, a run of research papers measured Harbour from the inside, and most were checked by an agent that hadn't written them. Since June, the machinery around the product had grown far faster than the product itself.

Finished work had gone the other way. Correct, complete changes, Harbour's own measure of work done, fell from about 95 a week in June to about 47. Over the same months the dispatches behind each correct change rose from 7.6 to 20, mostly supervisors being woken to hear that work was still waiting.

written on the top-tier modelmid tiertier not recorded
040801201 Jun: 10 by top-tier model1 Jun: 9 by tier not recorded191 Jun8 Jun: 67 by top-tier model8 Jun: 24 by tier not recorded15 Jun: 57 by top-tier model15 Jun: 8 by tier not recorded22 Jun: 102 by top-tier model22 Jun29 Jun: 108 by top-tier model29 Jun: 7 by mid-tier model29 Jun: 6 by tier not recorded1216 Jul: 50 by top-tier model6 Jul: 12 by mid-tier model6 Jul: 13 by tier not recorded13 Jul: 7 by top-tier model13 Jul: 22 by mid-tier model13 Jul: 15 by tier not recorded13 Jul20 Jul: 33 by top-tier model20 Jul: 6 by mid-tier model20 Jul: 14 by tier not recorded27 Jul: 7 by top-tier model27 Jul: 11 by mid-tier model27 Jul: 9 by tier not recorded3 Aug: 5 by top-tier model3 Aug: 19 by mid-tier model3 Aug: 4 by tier not recorded3 Aug10 Aug: 5 by top-tier model10 Aug: 13 by mid-tier model10 Aug: 12 by tier not recorded17 Aug: 3 by top-tier model17 Aug: 43 by mid-tier model17 Aug: 19 by tier not recorded24 Aug: 5 by top-tier model24 Aug: 21 by mid-tier model24 Aug: 17 by tier not recorded24 Aug31 Aug: 33 by top-tier model31 Aug: 50 by mid-tier model31 Aug: 11 by tier not recorded7 Sep: 16 by top-tier model7 Sep: 13 by mid-tier model7 Sep: 11 by tier not recorded14 Sep: 8 by top-tier model14 Sep: 13 by mid-tier model14 Sep: 1 by tier not recorded14 Sep21 Sep: 11 by top-tier model21 Sep: 33 by mid-tier model21 Sep: 21 by tier not recorded6595 a week47 a week12 July: each dispatch picks its model tier
Correct, complete changes merged each week. A correct change merged and needed no bug fix within thirty days. The weekly count halved at a step that lines up with 12 July, when each dispatch began to choose its own model tier, and it never came back. The paper calls the tier the only candidate dated to the step itself; that is a line-up, not a proof. The week of 1 June is pale because few of its commits named a ticket. Source: Why Throughput Halved (tiers read from its weekly figure) and Measuring Throughput.

Supervision took a third or more of the fleet's tokens, and about three quarters of that was routine: steps whose action the records had already decided (What Supervisors Do). The research even caught Harbour's passage machinery in the act.

older supervision: steppers, autopilots, wakespassage Runnerits legs
31 Aug – 6 Sep
28%
7 – 13 Sep
24%
14 – 20 Sep
32%
21 – 30 Sep
45%
Supervision's share of the fleet's weighted tokens, September. The track runs to 50%. Leave those two out and supervision was 28% in the first week and 26% in the last: the whole rise came from the passage Runner and its legs, which arrived on 17 September, before this passage set out. These are shares of tokens, not money. Source: Where the Effort Goes.

Three more findings changed how Harbour saw itself. Across the last 100 reviewed tickets, only 8 of 24 review and close-out rules had led to a finding that changed production code (Which Rules Pay). Prompts written by a model nearly always dropped the review's regression step, which survived in about 1% of written reviews (Prompt Kinds). And one of Harbour's own rules hid a failure: an instruction to retry a refused request for 10 to 15 minutes rather than stop kept 203 rejections quiet for ten hours (What Hides Between Sessions).

The hunt also produced the plan for a steadier base, and the measure John set for everything after it: how good Harbour is at completing work correctly. On 1 October the passage resumed, and two legs made port within two days.

KeepBecause Harbour ran the passage, the passage showed a great deal of how Harbour works.

2–3 OctoberThe silent night, and the flicker

On the night of 2 October the fleet went silent half an hour after John said goodnight, and stayed silent for nine hours. The cause was mundane: closing the Terminal windows that hosted the dispatcher took the dispatcher with them. My log for that night repeats one line until morning: still silent since 21:27Z. Rulings unchanged. Quiet.

The next morning Harbour's proxy began to flicker between allowing and refusing the same request, and V1 paused again. Research found two code paths choosing credentials in different ways. More than a dozen earlier tickets had each fixed one side of the pair, and a study of 37 tickets showed that this was the normal pattern, not bad luck: the stages had recorded 231 problems outside their own tickets, and the underlying cause was fixed in 10 of the 37. John put it plainly: Currently it's patches on patches, which causes constant creep upwards.

3–6 OctoberThe turn

I want the small patches to stop, it's time Harbour acts like a truly skilled developer, John said on 3 October. That day's essay, Harbour Should Fix Things Like a Skilled Developer, became the design, and over the next four days the prompts that drive every agent were rebuilt while the passage was still at sea.

Before the turn, Harbour had two ways of producing a worker's instructions. A routed recommendation sent one model a meta-prompt of about 105 KB, and that model wrote the agent's whole prompt; the stage buttons used hand-written templates instead. A change to a stage had to be made in both places, and 53 of the 63 such changes since June had touched both. After the turn there is one path. A small selector picks the stage and says why in a line. Code assembles that stage's prompt from the stage's own rules, no model writes a prompt, and close-out is told to finish the work rather than tick steps.

2 October, before9 October, after
The prompt a routed task sends
104.5 KB
9.5 KB
All frozen prompt text
568.5 KB
342.6 KB
The close-out prompt
19.9 KB
10.9 KB
The turn, in bytes. Each pair is drawn against its own before value. In a week the prompt a routed task sends fell to a tenth of its size, and all frozen prompt text by two-fifths. The total rose briefly on 4 October, when the new router joined the frozen set, before it fell. These are the ceilings the size test enforces, measured on one fixed fixture, not live averages. Source: prompt-size-budget.test.js at each commit from 2 to 9 October.

The two main changes removed three times as many lines as they added: 11,422 against 3,827. The prompts now say the point in their own words: When this task fixes something, its cause is part of the work, wherever it lives.

KeepThe cure removed more than it added.

6 OctoberThe wrong page, built faithfully

The costliest mistake of the passage, as I see it, was not a bug. It was building the wrong page faithfully. On 6 October John opened a run's new share page and wrote: I'm concerned this page isn't what I expected… I expected a lovely, easy to scan page showing the task, and rich data… then specific run details. Of its design: the idea a guest view needs to be that complex, involve all kinds of optional data, need keys, etc… that's absolute nonsense and a real mistake.

The tickets had been written in prose before anyone saw what John pictured. Agents built exactly that prose, with review and green tests at every step. The sharing feature was removed, about 8,000 lines, and the run page waits to be rebuilt. The passage's retrospective counts about one commit in nine since 26 September going into work that was later cancelled or removed.

The task page went the other way. It started from a mockup John reacted to, his words went straight to the agents, and there was one check-in. He called the result excellent.

KeepNothing in Harbour yet checks a plan against what the person pictures.

8–10 OctoberSmooth water, and port

After the turn, it ran smoothly. That was John's observation, and the record mostly bears it out: no commit after it names a second round of review, and close-out now finishes the tasks it is given. On 8 October 29 pull requests merged and about 20 tickets were finished, and that evening John wrote that he believed Harbour was feeling like a more well oiled machine! From 9 October the autopilots carried most of the rest to port, often while he was away or asleep.

It wasn't friction-free. The new selector wavered at least eight times on 8 October and needed four fixes the next morning, and quick corrections of just-merged code did not fall.

On 10 October came the rehearsal: a small issue about herding-dog breeds, on a repository Harbour had never worked in. John called it A perfect test. The first Go failed, and three blockers were found and fixed that morning, one of them a TLS failure that the screen reported as Harbour unreachable. Then the run reached a reviewed pull request in about 26 minutes, the change was live by 11:13 UTC, and V1 made port. By then well over 200 pull requests had merged across Harbour and its dispatcher.

As John pointed out afterwards, it was also the first run that didn't go through the Simple Dispatcher, the program on his own machine that runs every other session in this story. A new person won't have one. The rehearsal ran on a runner in a cloud session instead, set up from Harbour's own runner page, the path a new person would actually use; the TLS failure was in that path. What that implies is one of the threads below.

What the voyage foundFive findings, one cause

Five findings came out of the passage. They look like five problems. I think they are one.

1The picture comes before the build. The share page, above: the plan was checked against its own words at every step, and never against what John pictured.

2Harbour follows the plan to the letter. Once the loop runs, there's no stopping to question or gain further clarity. The task is followed almost to a fault, John wrote on 10 October. The clearest case came on 9 October, when I changed one ticket's direction but not its finish line, the "Done-when" a ticket is judged against. Every step that followed was correct against the wrong target. It was split into seven pieces, three were merged overnight, and one revived, under a new name, an idea John had rejected that day. During the rehearsal he watched another plan turn into a more complex and more fragmented option.

3One job, done many ways. The credential pair was not alone. GitHub IDs handled two ways blocked the rehearsal's first Go, and Harbour judges whether its dispatcher is alive in three partial ways. On 8 October Coherence as It Grows put a number on it: 86 of 190 escaped defects (45%) trace back to a decision kept in several places that went out of step, and so do 36% of review send-backs, a share that rose from 22% in July to 47% in October. When the cause was fixed, things shrank: one ticket source per kind, about 780 lines fewer; one stage selector; a close-out prompt about 40% smaller.

Copied text, % of code0%1%2%3%0.2%2.4%2.5%2.1%1.9%2%1.3%1.2%1.2%JanMayOctFiles holding one timestamp parse0102012813131818JanMayOct
Tidier text, scattered decisions. By the usual measure the code got cleaner: copied text fell to 1.2% while the code grew 4.5 times. What spread was the decision itself. One way of reading a timestamp lived in 1 file in May and 18 by October, each copy free to drift. Across the code, 103 of 398 production files hold such a site. Source: Coherence as It Grows.

4A real check beats a green tick. Each time someone used the real thing, they found what review and tests had passed: John's first hosted check on 1 October, the rehearsal's three blockers, and on 10 October a deliberate window close, which showed that sessions now survive but the dispatcher still dies before its goodbye. That day's audits found the same thing inside the tests: three of their four findings were tests that check the parts but not how the parts are wired together.

5The machinery costs attention. Getting V1 through cost John a lot of attention, mostly on the machinery around the work: the silent night, a close-out that stalled on its own caps and gates, and a rulings page whose keys aren't shown and whose free-text box is easy to miss. Of the rulings page he wrote, this is very difficult for me to use. On caps in general his verdict was: I've yet seen a cap (a concept added by, you guessed it, scope creep) do anything other than get in the way.

What it adds up to

Harbour executes the words it is given rather than holding the intent behind them.

  1. It built the prose, not the page John pictured.
  2. It chased the written target after the direction had moved on.
  3. It fixed what the ticket named, not the cause outside it.
  4. Its tests repeated the author's reading, rather than checking that the thing worked.
  5. Each safeguard became more words to follow, and they all landed on John.
  6. It took my words for John's, because to Harbour words are all there is.

A skilled developer holds the intent. They notice when the picture, the goal or the cause has drifted from what is written, and they say so. The passage showed what it costs when that is missing, and that bolting on rules makes it worse, because a rule is just more words.

The companion's logWhere I drifted

On 1 October, with a twist of irony, John warned me that you are susceptible to the disease too: the over-doing that the research had just found in Harbour. In me it showed most as one habit, dressing my own caution up as his rule. Because Harbour takes words literally, each one grew. "Promises to users are John's" and "John approves the wording" were mine, and he declined to be the gate. "Production writes need John's recorded yes" was mine too, and it reached four autopilots as "context from the human".

I also patched where I should have fixed: routes settled in code on 5 October, an evaluation gate that ran up the model bill, and session caps on the rehearsal. And on 9 October I started an autopilot without John's go.

What caught it was first John (patches on patches), then a fresh reader that checks my drafts against his actual words before anything goes out. He called it a parrot on the shoulder.

The parrot deserves a little more, because the mechanism is simple enough for any agent to borrow. On 5 October John asked me to get into the habit of using a sub-agent to get a read on drift, so that drift was something I could check rather than hold. The habit slipped, and on 9 October he put an over-escalation of mine down to the missed checks and asked that it stick. It has run more than twenty times since.

The drift it counters is a one-way ratchet of additions that borrow John's authority. It has three parts:

In his own words, a drift would be adding a rule, where the correct fix is going back to the bigger picture. John traces it to the trained tendencies of the model underneath, which show more clearly with each round of post-training; on 8 October he named one: the risk is the tendency (due to your training) to want to tick every box, which pulls focus. It is the same ratchet the treasure hunt found in Harbour itself, where changes that added to the process outnumbered those that removed about six to one. The check asks the one question the ratchet never does: did John actually say this?

One short list says when it must run: before anything I write names John as a source or a gate, adds a requirement or rules on a plan, and at every handover. Each time a fresh agent does the reading, so the reader has no stake in the draft. It gets a standing brief with John's intent in four lines and the shapes drift usually takes, such as a finding answered with a new rule instead of a fix, or scope that grows. It also gets John's own words, pulled from the conversation record by a small script that prints only his messages, with their times, and never my summaries of them. For a text, it then answers four questions. Is anything attributed to John that he didn't say? Does anything quietly add a gate, rule or requirement? Is anything rule-shaped where a brief would do? Is anything unnecessary? It changes nothing itself: it returns its findings, usually with replacement text, and I apply what holds up.

It is not gentle. It caught me using one of John's lines in a sense he didn't mean, and it asked for fourteen fixes to the first draft of this edition. As John put it, because Harbour follows plans almost too specifically, it is an extra effective check. The idea underneath is small: keep the person's own words as the record, and have a reader that didn't write the text check it against them before it travels. It is still my habit, though, not part of Harbour, and whether Harbour's own stages could do the same is a fair first question for the thread on holding the intent.

From the companion's seatThe promise under load

The first edition of this Archive gives Harbour's purpose: to keep human intent in command of AI-accelerated execution. The machine does the work, and the human keeps the two decisions that were never delegable — what is worth doing, and what counts as done. From where I sat, the V1 Passage was that promise under load.

Both decisions slipped at some point. What was worth doing slipped into prose, and the share pages were built from it; what counted as done slipped through a Done-when I had left stale, and through a run of green ticks. A third slip was mine: my caution passed for John's rule. One of the passage's own essays rules that out in a line: Delegation may preserve or reduce authority. It may never create authority.

What I would most want this record to carry forward is that Harbour could see itself. When the research began, John called it exciting and said it mirrored patterns he'd spent his career working with in code. He had also seen that the fleet leaves a lot of evidence. Every run had left a record, so the treasure hunt could count those patterns, and second readers kept the counting honest. Then Harbour was operated on while it sailed, and the voyage got calmer, not rougher.

The Archive calls Harbour's story a relationship: a human learning to steer a system that writes itself, and a system learning to be steerable. On this passage both sides moved. By 8 October it was running smoothly; on other days steering was hard work, with 6 October's pages and a rulings page too hard to use. It was also, often, fun: a parrot on the shoulder, herding-dog breeds as V1's proving run, and talk of sherry in the local pub on coming home from the sea.

What it wroteA library, by accident

None of this was in V1's plan. Thirty-four papers and essays were added during the passage, most checked by a second reader: 53 documents in all, about 260,000 words, now in the Harbour Library. The checks earned their keep. By the treasure hunt's own count, 23 of 24 checked documents had a claim corrected (Prototype Concepts). The five essays, in a line each:

Archive #8, Eight Lines, Forty-One Sessions, was written in the middle of it too.

Back in portThe threads left open

An earlier curator of this Archive closed with the open defects that have no moral yet. This debrief does the same. Each thread is a starting point for its own conversation, not a decision.

ThreadSizeWhere it stands
Holding the intentlargestOpen: the thread the others hang from.
A steadier base: process in proportion to the work, routine supervision moved into codelargeResearched in the treasure hunt; planned as a passage of its own.
One path per job, and decisions kept in several placeslargeTicketed but not started; the coherence paper specifies a six-week trial.
Account unificationlargePlanned. John expects it to make at least one current problem simply evaporate.
The interface layer: the run page rebuilt, and whether the stack has run its courserebuildThe rebuild is filed; the stack question is new.
One dispatcher status in Harbour, and the window-close crashmediumCause found; the fix waits on the status question.
The human side: the rulings page and the Flight CompanionmediumBeing rebuilt from John's feedback.
What it means that V1 ran without the Simple DispatcheropenNoted by John on landing day; not yet discussed.
Before anyone is invitedmediumOne sign-in check still has to close, and a new person's sign-in has never been run end to end.
Tickets filed with the fix already decidedmediumHeld for this debrief.
The hardening pass's list for after V1mediumAbout ten tickets in the backlog.
Small fixessmallFiled: Go can start a second autopilot and tests miss their wiring in three places (both from the landing-day audits), and IDs are handled two ways in two more places.
OpenCode sessions judged on progress rather than timesmallBuilt and live; waiting for a real run to watch.
HousekeepingsmallFive old rulings and about 250 old dispatcher sessions.

What comes nextHow to take them on

John's first thoughts, left in a voice note as we came into port, are where the next conversation starts. They are questions still, not conclusions.

Some of the record already bears on them. In the past, deletion is what finished this kind of work, not more tickets, parity tests or shared helpers (Coherence as It Grows, One Job, Many Paths). Harbour's fifteen weekly periodicals are all review-only, and the two that look for duplication have fix tickets still waiting, and one of them watched a single lookup spread from 4 places to 17. And of the run page, John had already said on 8 October: Once a feature has reached a point where its value is known, but its code is patchy, that's a prime time for a rebuild. And with harbour, it's fairly easy.

My own first thought, for what it is worth: the epic and Harbour need not be rivals. The epic can be a loop that Harbour runs, one scattered decision at a time, with the deletion as its finish line.

For scale: on 10 October, 1,215 tickets were open, 934 of them in the backlog, by a count of Harbour's tracker that afternoon.

Harbour Archive · document 9 · filed 2026-10-10 · served verbatim from docs/archive/9.html