Harbour Archive · Document 6 · A briefing
What Harbour is, how its work is organised, where the portfolio stands today, what the project's own instruments say about its health, and what a project manager would watch next.
Written 12 September 2026 by an outside Claude session handed a single-use read token, for a project manager meeting Harbour for the first time. Documents 1, 2 and 4 are museum editions; 3 is a project brief; 5 is an essay.
Grounded on: the live workspace API proxy at harbour.cat (read scope, 16:35–16:40 UTC, every issue paged), the public /kpis page (16:38 UTC), and the checked-out docs of both repositories (LinearViewer and simple-dispatcher). Figures marked instance-wide cover all 8 workspaces on the host, not only this one.
Harbour is a control plane for software delivery done by AI agents, run by one person. It reads a backlog from a task tracker, turns a ticket into a grounded prompt, dispatches that prompt to an agent session, and then checks the result against evidence rather than against the agent's own report. The landing page's phrase is "keep human intent in command of AI execution, one turn of the loop at a time."
It began in early 2026 as a read-only tree viewer for Linear. What compounded since then is not the model but the doctrine: staleness checks, a review-then-close-out split with a ledger of "what CI did not prove", a credential broker so no agent ever sees a live token, an operator decision queue, a north star the human alone may edit. The repository's CLAUDE.md alone is now roughly nineteen thousand words of that residue.
Two repositories make the product. LinearViewer is Harbour itself: the web app, the API agents call, the prompt engine, the observation and KPI surfaces. simple-dispatcher is the runner on the operator's Mac that polls Harbour's queue and drives each agent session through its phases. A third piece, Harbour OS, is the operator's in-browser workstation and is out of scope here.
It is dogfooded on itself. The tracker you are reading about is Harbour's own backlog, worked almost entirely by agents Harbour dispatched. One outward run on a foreign repository (28 August) proved the machinery ports but exposed that the non-Linear agent lane had never been exercised end to end.
The instance reports 12 users across 8 workspaces, but this workspace has one human: John Kershaw, the maintainer and operator. Everyone else on the board is an agent. The charter (a draft, dated 19 July, deliberately not yet adopted) says so plainly: "one maintainer and roughly one other user".
The division of labour is the design, not an accident of headcount.
Three agent roles sit above the workers. The Autopilot orchestrator chooses and accepts, never writes code. The Flight Companion is the observer: a chat surface that watches the fleet, narrates, and proposes follow-ups a human taps to approve; it is on rung one of a three-rung trust ladder (read-only, then supervised writes, then unattended). The Passage Runner flies a ratified passage leg by leg. All three answer to the same rule the handbook states: describe how things are going, but never redraw what "done" or "worth it" means.
Harbour's planning vocabulary is nautical and its scheduling unit is not a sprint. The terms a PM needs, in the order they nest:
LIN-nnnn, on one team. Six workflow states: Backlog, Todo, In Progress, Done, Canceled, Duplicate. No ticket carries a due date; the board measures flow, never schedule.front: names the front the work is on (twelve, from dispatcher-substrate to public-path); kind: names its origin: epic, research, decision, follow-up, or review-residue, the findings a review filed rather than fixed.main under docs/reviews/.docs/papers/. Seven exist; thirteen questions are queued. They are the project's process-health instrument./kpis estimates consumption from telemetry; the cap, not the backlog, is what binds.| Ticket | Passage | Ratified / proposed | Tasks | State |
|---|---|---|---|---|
| LIN-1876 | Close the loop, quiet the failures | ratified 3 Aug | 12 | ○Todo, never flown as written |
| LIN-2156 | Fix the lighthouse, then light it | ratified 20 Aug | 13 | ○Backlog (doubled as the runner's first acceptance flight) |
| LIN-2461 | Big Run, debt first then the deep night | ratified 2 Sep | 19 | ✓Done |
| LIN-2636 | Flight Companion V1, parity then trust | proposed 5 Sep | 14 | ✓Done, 12 of 14 made port |
| LIN-2730 | Flight Companion polish + OpenCode plumbing | proposed 10 Sep | 13 | ○Todo, awaiting John's yes; trim-safe minimum is legs A2 + B1 |
Every issue on the team, read this afternoon. Harbour's convention: ✓done, ◐in progress, ○not started.
Open work by project
1,004 open tickets. Scheduled means Todo or In Progress; parked means Backlog. Hover a bar for the split.
| Project | Todo | In Progress | Backlog | Done | Canceled + Dup |
|---|---|---|---|---|---|
| No project | 15 | 2 | 217 | 376 | 27 |
| Product | 58 | 2 | 84 | 450 | 57 |
| Autopilot, Recommendation & Prompt Engine | 41 | 2 | 97 | 194 | 19 |
| Simple Dispatcher | 26 | 0 | 93 | 152 | 14 |
| Dispatch & Execution Runtime | 30 | 3 | 55 | 84 | 12 |
| Quality, Periodicals & Measurement | 25 | 1 | 52 | 55 | 21 |
| Providers & API Unification | 12 | 1 | 62 | 104 | 12 |
| Platform Security, Robustness, Observability | 20 | 0 | 44 | 27 | 3 |
| UX, Theme | 6 | 2 | 32 | 58 | 3 |
| DevOps & Tooling | 4 | 0 | 10 | 57 | 6 |
| The Ship's Biscuit | 2 | 0 | 6 | 5 | 1 |
| Test Infrastructure: Local-Provider Migration | 0 | 0 | 0 | 28 | 0 |
Open work by front
Open tickets carrying each front: label. A ticket can carry more than one front, so the bars overlap and do not sum to 1,004.
kind:review-residue and 195 are kind:follow-up: 519 of 1,004. Reviews and close-outs file more than they fix, and what they file mostly waits (see the studies below).dispatcher-substrate leads with 249 open tickets, and the operating-model hub ranks finishing the Linux dispatcher host above every other phase because "the September incident set is almost entirely the Mac host".| Ticket | Project | Priority | Title |
|---|---|---|---|
| LIN-751 | Autopilot & Prompt Engine | High | Realtime chat interface for work in flight (the Flight Companion epic; 20 open children; open 67+ days) |
| LIN-2114 | Dispatch & Execution Runtime | High | Move observation-type sessions out of Claude Code into a simpler cloud harness (epic) |
| LIN-2246 | Providers & API Unification | High | Decompose the 4,388-line workspace-api route file by URL group |
| LIN-1933 | Quality & Periodicals | High | Periodicals: target-repo selection at dispatch (active 25+ days) |
| LIN-2634 | Autopilot & Prompt Engine | Medium | Rung-two evidence: five paired mornings, agent-graded, with minutes-to-first-decision |
| LIN-2459 | No project | Medium | Workspace prep copies the source checkout's untracked settings into every clone |
| LIN-2811 | UX, Theme | Medium | Gate the chat reveal on a shared pinned-to-bottom predicate |
| LIN-2808 | UX, Theme | Medium | Preserve scroll position and indicate new conversation updates |
| LIN-2089 | Product | Medium | Ship Journey: waypoint dot is six times the step length |
| LIN-1675 | Product | Low | Animated ship journey map from roadmap data (epic; the north-star reading names it as drift) |
| LIN-2422 | Dispatch & Execution Runtime | None · Blocked | Phase 1 guest provisioning script for the Linux dispatcher box |
| LIN-1785 | Dispatch & Execution Runtime | None | Provisioning script for the Linux dispatcher box |
| LIN-2397 | No project | None | Factor the shared GitHub App auth flow out of two route files |
"In Progress" is a loose signal here. The team's own practice, recorded on LIN-1981, is to move a ticket nobody is actively working back to Todo, so a long-lived In Progress row usually means an epic or a stalled item rather than a live session.
Harbour publishes its own operating figures on a public, unlinked /kpis page. These are instance-wide, thirty-day windows unless stated.
The Recent Headwinds review counts mainline commit units per active day on the Harbour repo. It refuses to give a single number, and so should you: one 52-unit lane-run day on 23 August lifts the latest window's average from below the prior baseline to well above it.
| Window | Span | Units | Per active day |
|---|---|---|---|
| A · 8 Jun – 6 Jul | 28 d | 459 | 16.4 |
| B · 6 Jul – 3 Aug | 28 d | 230 | 8.5 |
| C · 3 – 23 Aug | 20 d | 117 | 8.4 |
| D · 23 – 29 Aug | 7 d | 97 | 13.9 (7.5 without the lane day) |
The 23 August lane run itself is the best single day on record: fifteen long-lived sessions, 52 tickets to Done, 43 PRs merged, zero faked closes, wound down deliberately at 71% of the weekly budget. The method it codified (one session carries an ordered ticket list with a file carve, without pausing) is now one of three ways work is flown, alongside per-step autopilot and per-leg passage runs.
A routing proposal dated 11 September priced 1,068 worker sessions across 155 issues over nineteen days at $6,687 API-equivalent. The orchestrator tier is the single largest line; implementation is where the cheaper model already runs.
| Kind (model) | Sessions | Share of spend | Median cost | Median minutes |
|---|---|---|---|---|
| autopilot orchestrator (Opus) | 113 | 23.8% | $9.27 | 140 |
| implementation (Sonnet) | 172 | 19.1% | $4.54 | 19 |
| review (Opus) | 188 | 14.0% | $4.67 | 11 |
| close-out (Opus) | 156 | 11.2% | $4.17 | 10 |
| research (Opus) | 78 | 8.8% | $6.37 | 14 |
| plan-review (Opus) | 118 | 8.1% | $4.45 | 10 |
| plan (Sonnet) | 127 | 3.8% | $1.74 | 7 |
| implementation (Opus) | 12 | 4.1% | $18.72 | 57 |
The evidence behind the routing: Sonnet implementations passed review first time in 84% of 115 cases at a quarter of Opus's median cost, and terminal failures are rare on every kind. The proposal's whole saving is effort, not model: three effort reductions (research, review, close-out) touch a third of worker spend and would save roughly 10–15% if published curves hold. Nothing has been changed yet.
The capacity model in one sentence. The fleet runs on one prepaid subscription that buys thirty to sixty-five times its price in list-rate compute when the week has headroom; two of four measured weeks hit the cap, and the essay "The Cheap Ships" (5 September) sets out five levers, in compounding order, for running at full capacity and eventually stepping off the subscription: an observer that does not think, a generated wiki instead of a research phase, a cheaper implementer (Dash), routing that keeps judgement expensive on purpose, and context discipline.
Harbour measures itself. Seven papers landed between 5 and 12 September, each answering one question about the process from the tracker's own data. Read together they describe a system whose gates work but whose exhaust is piling up.
| Question | Finding | What it means for a PM |
|---|---|---|
| Why does a plan go round plan-review more than once? | 83 of 94 plans were sent back; 33 went round twice or more. Every send-back asked for one more member of a list the plan had already built, never a different design. The loops cost a third of all session time. | The plan gate is a list-completeness gate, not a design gate. The rule changed on 12 September to "argue the class, not the member"; the re-measurement is queued for 26 September. |
| How do tasks generate tasks? | 2.1 tickets created per one closed over sixty days, above 1.3 every week. A generated ticket that is worked produces 1.12 more. Half of all filings are never worked; close-out follow-ups and review findings are 63% of that pile. | Intake outruns closure by design of the pipeline itself. Backlog growth is a process property, not a demand signal. |
| Does the system run out of tasks? | Collapsing breakdown trees to one unit, each unit causes 0.74 further units. Three in four cause none; a unit that reaches Done causes 1.52. A tenth of units carry 73% of all causation. | Worked work is above replacement. The population number sits under one only because most filings are never touched. |
| What is in the never-worked pile? | Of 40 filings read by hand, 34 are still true at HEAD and 38 are actionable on their own. 23 of 40 are the parent ticket's own unfinished scope, filed next door. | The pile is real work in the wrong place, not stale noise. It cannot be closed by age. |
| Does a close-out say so when it files its own scope? | Always: 23 of 23 name the filing. But filing was the mechanism by which the ledger item discharged and the parent closed. | Led to John's 12 September ruling: an inside-scope item can no longer be discharged by filing a ticket. |
| Does the writing get longer faster than the ideas? | Yes, by about 1.6×. Reviews tripled in length in three months while findings doubled. CLAUDE.md grew 11.5× while its section count fell. | Reading cost is a real tax on the one human. The comprehension-debt review that should have caught it never fired. |
| What cost levers are already written down? | Thirty-seven, across five stages of a ticket's life. Most measured once, on one day; none measured past the gate that judges its output. | A map for choosing the next experiment, not a ranking. Cost work is rich in hypotheses and thin on controlled reads. |
The north star, v2, is titled "the self-funding loop" and makes one economic claim with seven clauses behind it:
Harbour completes verified backlog work at a cost and cadence a solo operator can sustain, funds itself doing it, and proves every word.
The roadmap report of 2 September scores the portfolio against that text. Its verdict, condensed: the operational core (evidence integrity, bounded runs, spend attribution, silent-failure work) is aligned. The drift it names is elective interface work: UX & Theme, The Ship's Biscuit, the Ship Journey map, and Flight Companion expansion whose effect on operator minutes has not been demonstrated. Two unscorables sit underneath: "verified task" is still a proxy (the KPI counts terminal markers, not discharged ledgers, tracked as LIN-1878), and the first published forecast (LIN-1626) is unstarted while blocking six tasks.
The report ends with the question it wants the human to answer: should supervised conversational control count as a north-star outcome in itself, or only when it demonstrably reduces operator minutes and sessions per verified task? Nine days on, that question is still open, and the Flight Companion has had two passages since.
| Programme | Anchor | Shape | Status |
|---|---|---|---|
| Operating model path | LIN-2569 | 19 tickets in five phases: an adaptation memo, secret scans and deploy witnesses, a declared invariants registry with a deterministic Measure job, an opt-in estate report, then sensors. Ranks the Linux dispatcher host above all of Phase 3. | ○Phases 0–2 Todo, 3–4 Backlog; filed 9 Sep |
| Path P0–P4 (self-funding) | LIN-1625…LIN-1652 | 30 tickets from cost-per-verified-task through forecasting, threat model, public projection, hosted tenancy with a ring-fenced free tier, to governance and an AGPL release. | ✓P0.1 done; ○29 Todo |
| The Cheap Ships | LIN-2686 | Fold the five cost levers into the run. 8 open children. | ○Todo, High |
| Linux dispatcher host | LIN-1781 / LIN-1785 / LIN-2422 | Run the runner headless under tmux, off the Mac. The tmux driver exists; provisioning is the blocked piece. | ◐In Progress, one child Blocked |
| Flight Companion | LIN-751 | The observer chat, rung one of three. V1 shipped 5–10 Sep; polish passage awaits a yes; rungs two and three gated on John. | ◐In Progress since July, 20 open children |
| Operator decision queue | LIN-1721 | The rulings inbox and its escalation lifecycle. 41 children, the largest single generator of new work on the board. | ○Todo, 10 open children |
| Trivial onboarding | LIN-614 | Account → Connection → Workspace; add any number of sources from one connection. The path to "a stranger can trust it in one sitting". | ○Todo, 2 children |
Drawn from the last Recent Headwinds review (29 August), the Drift & Coherence review (29 August), the incident file, and today's live reads. Severity is the reviews' own.
| Risk | Evidence | State today |
|---|---|---|
| critical The review layer's own cadence instrument is unreliable | 10 of 15 periodicals read "never" on 29 August, including four whose reports had landed; the ledger keyed on a dispatch being taken, not on a report landing. | Partly fixed (LIN-2385). Live today: 7 of 15 due at 13–14 days, 8 still "never" in the 30-day window. The four never-run corrective reviews include security and dependency supply chain. |
| critical → fixed The direction layer served the wrong north star | The proxy served v1 while the doc was v2; the alignment reading went stale, then empty. | Fixed. Today the reading is v2 and fresh at 9 days, with a doc-hash drift check (LIN-2254). |
| high Single host, single operator | The runner lives on one Mac; the watcher deliberately does not relaunch a dead dispatcher (a merge to main is the only remote lever). The September incident set is "almost entirely the Mac host". | Open. The Linux host is In Progress with its provisioning child Blocked. |
| high Fix-induced chain in the runner's wake/resume path | LIN-2297's root-cause fix did not bound its class; LIN-2331's own title says so. Seventeen stall/refire tickets in one week. | LIN-2331 still Backlog. The OpenCode "hold survives the runner dying" leg (B1 of the pending passage) is the next move on this surface. |
| high Gate deferrals expiring unwatched | LIN-1661 and LIN-1873 carried "not before 25 August" in their titles and passed the date with no action and no fresh ruling. | LIN-1661 Todo, LIN-1873 Backlog. Nothing on the board watches a title-encoded date. |
| high Cost figures on a hand-edited price table | Sonnet 5 introductory pricing ended 31 August; 36% of the window's spend rode on it; the pricing table has no expiry field. | Unverified whether the table was updated. Every cost figure in this brief inherits that uncertainty. |
| medium Latent import cycles | Three cycles among the KPI and budget modules, five more at the provider seam; all latent, none live. The provider-auth inversion has worsened four reviews running; its fix (LIN-675) sits in Backlog. | Open. |
| medium The agent lane on non-Linear backends was never exercised | The outward run on 28 August found five of thirteen natural agent reads return 500 on GitHub, and every generated prompt told the worker to update "Linear". | Two of eight findings Done, rest in Backlog. |
| medium No schedule signal at all | 0 of 2,769 issues carry a due date. Every timeliness reading is flow health, never schedule health. | By design. Passages are the only dated commitments, and they are dated by ratification, not by landing. |
| medium Reading load on the one human | Reviews tripled in length; comment words per issue grew 3.6×; CLAUDE.md is 19,000 words. | Named in a paper; no corrective ticket found. |
The one incident file worth reading first. On 8–9 August every Linear call from the proxy returned 401 for twelve hours. No data was lost; three sessions detected it and parked cleanly. The root cause was never found, an amplifier was, and the fix was held unmerged pending evidence. It is a model of how this project writes incidents: times with their clock provenance, a corrections section, and a monitor named for what it could not prove. Whether that monitor ever fired is itself a Todo ticket (LIN-2577).
These follow from the evidence above and respect the line the project draws: a PM here decides what is worth doing and what counts as done, and leaves the how to the loop. Each is one decision, most of them John's.
/workspace/:key/observation): every agent session, live, with a Rulings tab for decisions waiting on you./stack?view=digest: the ranked frontier with why each item ranks where it does./north-star: the live intent plus the freshness-gated alignment reading./periodicals: which weekly reviews are due, recent, or never ran./rulings: every unanswered decision, with the option to propose an answer (never take it)./issues/{id}/cost: dollar cost of one ticket, refusing to total partial data./issues/{id}/brief: a present-tense distillation of a ticket that supersedes stale wording.Counts by state, project, priority and label were computed from all 2,769 issues paged over the proxy today. Passage and epic detail came from the same API. Throughput, cost breakdowns and risk severities are quoted from the project's own reviews and papers at the checked-out HEAD, and each is dated in the text so you can judge its age.
What this brief could not see: the issue list carries no creation dates, so ticket ages here are the reviews' numbers, not mine; the repository checkout was shallow, so commit velocity is quoted rather than recomputed; the KPI page aggregates all eight workspaces on the host; and cost figures are API-equivalent list rates on a subscription, so the cash cost is lower and the "verified" in cost per verified task is still, by the project's own admission, a proxy.
Harbour Archive · document 6 · filed 2026-09-12 · read scope, every call audit-logged · served verbatim from docs/archive/6.html