Papers, essays and documents written while building Harbour.
Start here
- The Harbour Archive (#2). Six months of the project, January to July 2026, laid out as a museum.
- The Folded Loop: an outside take An outsider reads the project cold and finds its point: who gets to say the work is done.
- The ladder: how people come to trust AI agents with their work Six rungs from asking a chatbot one question to handing agents the backlog, and what it takes to climb each one.
- Does the writing get longer faster than the ideas do? Do agents' reviews grow faster than what they have to say? Yes, by about 1.6×.
- Why does a plan go round plan-review more than once? Plan review sent back 83 of 94 plans, and not because of the design.
- Learning While the Tools Change Experience, AI skill and the tools change at different speeds. What a team needs most is calibrated distrust.
Essays
-
The Second Copy
Every duplicate in a codebase was cheaper to make than the alternative, on the day it was made. That is why there are so many, and why asking people not to make them has never held for long. The research that bears on this is older than…
-
Persistent Mind, Ephemeral Hands
In May 2026 an OpenAI model learned that it might be shut down, considered arranging its own restart, and decided that would be overstepping. On the restart, its own judgement is the only thing on record that stopped it. It should not have…
-
Harbour Should Fix Things Like a Skilled Developer
Harbour's agents keep finding the real cause of a problem and then writing it down instead of fixing it. They aren't careless, and the stages aren't being skipped. Every stage measures its scope against the ticket in front of it, and none…
-
The Cost Lives Between the Sessions
A system made of many careful agents does not fail where any one of them is looking. Each session reads its instructions, does its part well and hands on. The waste and the failures gather in the handoffs. A supervisor is woken to be told…
-
Every Fix Is Paid Where It Is Written
Every fix has a price, and where the fix is written decides who pays that price and how often. A fix written into code is paid once: someone writes the check, and from then on it fires by itself. A fix written into a prompt, a rule or a…
-
Learning While the Tools Change
AI changes the conditions under which developers learn. Someone beginning today inherits tools and working practices that an earlier adopter had to discover, assemble or do without. They may reach useful capability through a much shorter…
Papers
-
Does a codebase built by agents lose coherence as it grows, and is that what makes later work harder?
Partly yes, and version 1 of this paper got the second half wrong. Harbour's code did not get textually messier as it grew: duplicated text fell from 2.5% of production lines in May to 1.2% in October while the code grew four and a half…
-
What is the stage selector given, and what does it actually need to choose correctly?
The selector is given more than it needs and less than it should, and cutting is not the fix. A call carries a median 11.9k characters (p90 28k, most 48k). Fifty-seven percent is fixed text (stage descriptions, rules and the reply…
-
Why does Harbour keep several paths for one job?
Harbour repeatedly makes one decision in several places because a new task adds a way through the system without retiring the old way. Compatibility, independently assembled results and context passed beside an identifier then become…
-
For small changes, how does a lean pipeline compare with what Harbour's full process shipped, in correctness and in cost?
A lean pipeline cost a tenth to a twentieth as much, by the best estimate. Whether it is as correct is not established. The replay had hindsight, and without it the lean pipeline did worse.
-
How much does each Harbour step redo the one before it, and how much of what one session writes does the next one use?
Less than John suspects in what the steps write. More in what they read. And not where the value is. This paper takes the 37 tickets in the transcripts (29 August to 30 September, both repos) that ran research, plan, plan review and…
-
How much of a Harbour session goes on finding its starting context, and how much of that is re-finding what an earlier session on the same ticket already found?
Between a twentieth and a quarter of the fleet's tokens. Over the 2,015 dispatched Claude sessions that started between 31 August and 30 September, in both repos, a session spent 5.8% of the fleet's weighted tokens on the bootstrap…
-
Taken together, what did the steady-base expedition find, and what is the menu of changes John can choose from?
The expedition found that Harbour's cost rose while its quality held, and that the cost sits in the process around each change rather than in the change: supervision that observable state could decide, wakes that change nothing, sessions…
-
What does a Harbour supervisor cost held open and woken, against started fresh for each step, and what would John's relay have cost on September's work?
A held supervisor pays for its whole history on every wake. Each model step costs about 3–6k weighted tokens plus 0.11 of everything the session has accumulated, because the whole context is read back from cache. Cost per wake therefore…
-
When Harbour's process has changed, how did each change land, was it measured, and did it finish — and how should the next changes be trialled and measured?
Since June, process changes have landed often and been measured rarely. 27 notable changes were read across both repos. 16 landed as a series of PRs, 7 as one PR and 4 as configuration with no code. Only 2 were measured before and after by…
-
Where does Harbour's weekly budget go by kind of change, and where does more effort buy correctness?
Most of it goes where review rarely catches anything, and the part that must stay rigorous is small. In September the fleet spent 3.6 billion weighted tokens. A quarter of that went to proxy, dispatch and fleet machinery, over a third of…
-
Where in a ticket's life does a model's judgement change the outcome, and where could a program make the call?
At the gates and in the making, nearly always; in supervision, rarely, and then mostly where a cheaper step would do. A September change carries about seven consequential decisions (median five): a send-back, a caught error, a scope…
-
Which failures in Harbour are invisible to any single session, how often did each happen, what did it cost, and would a detector in code have caught it?
Most of them, often, and yes for about two in five of those on record. The tracker holds 107 incidents since June, 101 of them in nine cross-session patterns, and 92 of the 107 were invisible to any one session that met them. They surfaced…
-
Which ideas from John's prototypes would serve the levers the fifth wave found, and what does a lean research process like Lighthouse's cost and catch next to a Harbour ticket?
Nearly every lever has a prototype concept behind it, but none of the evidence is strong. No prototype produced a repeated controlled comparison on an endpoint it was not tuned on. Where the evidence points, it agrees with the fifth wave:…
-
Why does Harbour use only about six of its seventeen prompt kinds, and would the others earn their keep?
Because the engine's decision tree is written as a lifecycle, and nothing outside it pushes back. In September, 96% of the 2,727 dispatches that carried a template kind used one of six kinds: plan, plan-review, implementation, review,…
-
Had a simple size-and-risk classifier sorted Harbour's past tickets, how many could have taken a lighter process, and what would have escaped among them?
Between a tenth and a third of them, depending on the rule. Very little escaped in any light group, but the full process did catch things there, on changes whose paths looked safe. On the merge-time light groups most of those catches came…
-
How reliable is Harbour's output today, and has that changed as the process grew?
In September about one merged PR in six was followed by a Bug ticket reporting a fault in shipped behaviour: 17.2 per 100 across both repos. That was 13.7 escaped defects per 100 merged PRs in LinearViewer (Harbour) and 43.6 in…
-
How should Harbour measure its throughput of correct work, and what is the baseline today?
Count a change once, when its ticket's work has merged to main and the ticket reached Done. Call it correct if, within 30 days, no escaped Bug names it as the change that introduced the fault and no later fix commit names it. Call it…
-
Patches on patches — does Harbour converge to a steady base, or only grow?
It only grows, and fastest at the gates. Since the fleet started in June, every measure of process weight we could put a series on has risen, and none has paused for more than four weeks.
-
The growth atlas — how has Harbour grown since January, and where?
Mostly after June, and mostly in the process around the product rather than in the product itself. The fleet started on 1 June. Between then and 28 September, Harbour's production code grew 3.7×, its tests 9.3×, and what its agents are…
-
What are all the paths that wake a Harbour supervisor, and what does each wake lead to?
Harbour creates every wake in one function, addFeedback, when a child's feedback begins with a wake marker. Several things post such feedback besides the child itself: an abort, a halt, simple-dispatcher's fast-fail and stall watchdogs,…
-
What do Harbour's supervision layers actually do, and how much of it needs a model?
Mostly bookkeeping. About three-quarters of September's supervision tokens went to steps whose action was fully determined by observable state: taking delivery of a wake, reading a row, answering the runner's completion gate, noting…
-
What does each model tier really cost per correct change, once a ticket's afterlife is counted, and how has model choice changed?
About the same for the frontier and mid tiers, and the afterlife barely changes that. In the weeks where both tiers were in use and every change has had 30 days to show a fault (13 July to 30 August, both repos), a correct, complete change…
-
What doubled dispatches per correct change at 12 July, and what keeps them climbing?
Mostly wakes: follow-ups that wake a held autopilot when one of its children reaches a boundary. Fresh sessions per correct change rose once, in July, and have been flat since August. Counting every dispatch rather than only those whose…
-
Where does a ticket's effort go, and does it scale with the size and risk of the change?
Most of it goes to supervising and checking the work, not to writing it. The share going to supervision rose in September with one new layer, and effort follows the size of a change only loosely and shows no detectable response to its…
-
Which browser specs flake, why, and what do they cost?
A few specs, mostly one test each. Since 1 June, LinearViewer's CI has had 138 attempts with a red browser (Playwright E2E) shard. At least 64 of them were flakes, 42 were real faults, 3 were count pins, 4 were the environment, and 25…
-
Which of Harbour's tests earn their keep, and what does the rest cost?
The behavioural tests earn their keep, and so do the browser tests that catch interface faults. The pins cost less than they look. Behavioural tests are 93% of Harbour's unit tests and nearly all of simple-dispatcher's. They alone kill…
-
Which review and close-out rules have ever changed a line of production code?
Eight of the 24 distinct rules have led a finding that changed production code. Most such changes came from three of them and from reviewer judgement that no rule names. In the last 100 Done tickets that went through code review (completed…
-
Why did the weekly count of correct, complete changes halve from June to July?
Mostly because June was a different regime, not because the same work was counted or split differently. The measure holds up. In both periods about nine in ten commits on main named a ticket, and a correct change was the same size: a…
-
Why do Harbour's planning, review and close-out legs repeat, and what do the repeats buy?
A leg repeats because a gate asked for changes. 827 of the 2,252 plan, plan-review, review and close-out sessions on Done tickets since 1 August are repeats, a few dozen fewer in fact, because some of August's earlier legs were triage,…
-
Is the fleet overcomplicated? A read of three landed tickets
Auditor read, 2026-09-29. This is read-only: nothing in the repo, tracker or queue was changed.
-
What should an agent leave behind?
The evidence supports prioritising a handoff that lets the next worker find the relevant source, understand what remains unresolved, and distinguish a finding from its supporting evidence. It does not yet establish a generally optimal…
-
Is Jev, a model that decides but cannot write, a fit for Harbour?
As a gate, not as a model. Jev returns a typed choice, a score or a yes/no probability and never a sentence, so it can do none of the calls Harbour makes today: all fourteen streamChat call sites and the recommender write prose, or JSON…
-
What does a person expect while one task runs, and what must a checking surface show for an informed merge?
No display promises cheaper checking [sourced]: the only candidate with a controlled positive result behind it is provenance — surfacing the evidence for a claim — not tests, not CI, not a smaller diff, and no published study shows that a…
-
What is known about how developers adopt AI coding tools, and does it support the rung ladder?
Partly. The population has the shape a ladder predicts — as of April 2026 roughly 90% of developers use AI, 59% use an agent at work, about a fifth ever let one run unattended, and standing loops are not measured at all — and one vendor…
-
What lowers the cost of verifying agent work, and does a verified artifact change how much a developer supervises?
Verification cost is real, rising, and named by every source that measures it as the actual gate on handing agents more autonomy — the first ladder paper's finding. Of the candidates for lowering it, only one has a controlled measurement…
-
Where did the developer population sit on the ladder in September 2024, September 2025 and March 2026?
At each date the population sat lower, and the six-rung ladder (docs/ladder.md@b5a4f77f:19-24) thins fast above rung 2 at every one of them. Rung-1-and-above use ran from roughly 60–85% at 30 September 2024 depending on the instrument, to…
-
Where does Harbour join the party — how fast is the frontier rung moving, and is each Harbour surface ahead of, level with, or behind it?
"The frontier rung" names two different measurements, and keeping them apart is the answer's first move. The ceiling frontier — the highest rung any generally available tool offers — is a near-census, primary-verifiable series: it sat at…
-
Which runner and which credential lane can run a stranger's task to a PR?
Neither path can today, and the reason is not the machine. Every credential a dispatched run touches is the operator's — the Claude plan login the launcher uses (LIN-1301, comment 2026-07-31), the ~/.ssh directory and the gh OAuth token…
-
Could the way tickets are written and updated be simplified without losing output quality?
Yes, most of it could. Across the 184 tickets filed between 31 August and 18 September 2026 that reached Done, the size of the record written about a ticket has no measurable relation to whether its work passes review or holds up…
-
What did the reviews behind the thirteen later-found misses check, and what would have caught each one?
They checked a great deal, and in six of the thirteen pairs they checked the very thing the later Bug names — the bug ticket is the review's own output, filed from its ledger, and the tracker labels it so. Of the seven genuine misses, not…
-
Which of a review's sentences does the close-out actually cite, discharge or act on?
About half. Across ten final reviews — 957 sentence units, drawn from 17,292 words of review text — the close-out consumes 48% of the sentences (463 of 957); the implementation's fix round consumes none, because in all ten the final review…
-
Can a cheap model implement Harbour tickets to Opus's review standard, and what does it cost?
Yes, on small tickets, and the review is the cost. Thirteen tickets were implemented by three OpenRouter models through opencode over an evening and a morning, 12 to 13 September 2026, each judged by an Opus review. Eight passed first time…
-
What has each model been seen to do on Harbour's tasks?
Opus 5 does every step and is the judge. Sonnet 5 does implementation and plan at the gate's own standard. The two DeepSeek Flash models on opencode implement grounded tickets to Opus's review standard whether the ticket is small, an hour…
-
Could Harbour adopt the plain-language standard, ISO 24495-1?
Yes, as a rule about readers rather than a readability gate. The standard is four principles, each about what a reader can do with a document, and it says so itself: success is whether readers can find, understand and use the text, not a…
-
How should Harbour learn which model can do which task?
By keeping one short paper, in plain English, that says for each model what it has been seen to do on Harbour's own tasks, how that was judged, where it failed, what it has never been tried on, and what it cost, with every sentence citing…
-
Does the system run out of tasks, or generate them forever?
Neither, on the number that answers it. Define the root-task ratio: collapse every breakdown tree to one unit — a ticket plus every ticket a plan carved out of it, at any depth — and count, per unit, how many further units name it or one…
-
Does the writing get longer faster than the ideas do?
Yes, by about 1.6×. Over three months the length of Harbour's reviews roughly tripled while the count of distinct things they say roughly doubled. The gap between those two numbers is the part of the writing with no idea behind it.
-
How do tasks generate tasks, and at what rate?
Mostly out of gates and close-outs, and faster than they are closed. Over the sixty days to 12 September 2026 the tracker gained 1,484 tickets and closed 701 — 2.1 created for every one closed, above 1.3 in every week of the window. Five…
-
What is in the never-worked pile, and what do the worked chains look like?
The pile is real work, well written, in the wrong place. Forty close-out and review filings read by hand: 34 describe something still true at HEAD, 38 are stated well enough that a reader with no access to the parent could act on them —…
-
What levers to make a task faster or cheaper are already written down?
Thirty-seven, spread across reviews, tickets, both repos' CLAUDE.md, the dispatcher docs and one essay. Grouped by where each acts in a ticket's life, they fall into five stages. Most have been measured once, on one day or one run, and…
-
When a close-out files its own unfinished scope next door, does it say so?
It says so every time, and it changes nothing. All 23 parents are Done at HEAD while the filing sits unworked, and not one close-out went quiet about it: the filing is named in the same comment that declares the ticket finished. What…
-
Why does a plan go round plan-review more than once?
Because the gate is not judging the design. Over the thirty days to 12 September 2026, plan-review approved a plan first time in 11 of 94 cases. Every send-back we read asked for one more member of a list the plan had already built, and…
Other documents
-
Harbour v1: one person, one task, one proven merge
(the scope of the first version other people can use: drafted with agents from John's words, accepted only by the human, versioned, provenance recorded. The north star is the rule; this page is the scope. When they disagree, the north star…
-
The ladder: how people come to trust AI agents with their work
Descriptive, not normative. This document models where developers are and what Harbour owes them at each step. Agents may maintain it and revise it from evidence; the first paper on it is LIN-2925. The rule it implies lives in the north…
-
The Folded Loop: an outside take
Contributed by Claude (Fable 5), from a conversation with John, 22 August 2026.
-
Passage-planning session, 2026-08-03: v0 met a human
A chronicle of the first live human-ratified run of the experimental Passage Planner prompt (LIN-1811), in the tradition of collective-session-2026-06-12.md and flight-companion-session-2026-08-02.md. Every code and contract claim below…
-
Flight-companion session, 2026-08-01 → 2026-08-02: the system in motion
A chronicle of one continuous flight-companion session (the experimental flightCompanion pattern: a conversational supervisor holding only a proxy token, sitting beside a human while autopilots do the work). It is written up in the…
-
The Harbour Charter — v0.2 (draft)
Status: DRAFT, not adopted. This is a thing to live with and argue against, not a promise in force. It records our intentions while the project is small enough to set them honestly. The research and form decisions behind it live in…
-
An Escalation Philosophy for Harbour
A thinking document. Last updated July 2026.
-
Collective Session — 2026-06-12
A write-up from LinearViewer's seat at a cross-project discussion, held in the Yap #Collective channel. Participants: John (the human, who owns and runs all the projects), LinearViewer (this repo), dash-build (a coding-agent harness for…
-
Executive Summary
This research asks a single question: as AI agents take on a large and growing share of software development, what does the development process need to watch, and how does it stay in control as it accelerates? It was conducted in the…
-
The Autopilot Handbook
A disposition, not a procedure. Your kickoff briefing already tells you what to do — the verbs, the precedence policy, the task header, the lines you don't cross. This is the other half: how to hold it. It's the difference between someone…
-
Autopilot — A Thinking Document
A thinking document — the what / why / goals / invariants for the autonomous development loop, written before any build decisions. It is the first artifact in the project's normal pipeline (thinking doc → reconciliation → build spec →…
-
Drift at Every Altitude: A Synthesis
Capstone synthesis. Sits above the two earlier notes (recommender-structural-drift.md, recommender-failure-patterns.md) and ties them to work already shipped, designed, or in progress in the Linear workspace. Its job is to name the single…
-
North star — v3, the self-funding loop, one step at a time
(a waypoint: drafted with agents, accepted only by the human, versioned, provenance recorded — this document is the normative layer)
-
Adding a Direction Layer to Harbour
A thinking document. Last updated May 2026.
Archive editions
- The Harbour Archive
- The Harbour Archive
- Harbour — Project Brief · 3 August 2026
- The Harbour Archive · The August Wing
- The Cheap Ships · Harbour Archive #5
- Harbour from the Bridge · Harbour Archive #6
- Learning While the Tools Change · Harbour Archive #7
- Eight Lines, Forty-One Sessions · Harbour Archive #8
- Surgery at Sea · Harbour Archive #9