Harbour project brief
generated 3 August 2026 by Claude

The bottleneck moved.

Harbour is generating work as fast as ever and converting a third less of it into shipped code. That gap is the story.

period trailing 30 days scope 2 repositories · 1,810 tasks · 14 arcs upstream report 0d old · fresh
◐ AMBER Delivery is down by a third. The effort went into verification — and the instrument that should have caught it stopped running.

Tickets reaching code fell 39% (115.2→69.8/wk) and mainline delivery units fell 37%, while ticket creation held flat at about 190/wk — so conversion, not capacity, is what changed. Work per ticket is flat (+4%). Production code fell 54% while test code held steady. 1 workstream stalled, 4 under pressure, 5 healthy. No schedule risk is measurable: 0 of 1,810 tasks carry a due date.

Tickets created 193→190 per week · flat
capacity to generate work is unchanged
Reaching code 60%→37% share of tickets a commit ever cites
this is what fell

What this is

source repo + CLAUDE.md

Harbour is a provider-agnostic control plane for AI-augmented software development. It reads an issue backend, generates a grounded prompt for a specific task, dispatches that prompt to an AI agent, and verifies the result against real evidence. It began as a read-only Linear project tree and grew into the cockpit for that loop.

The system is two repositories. Harbour is the web control plane — provider adapters for Linear, GitHub Issues, GitHub Projects and a local store; a dual-path prompt engine; a dispatch queue; and a source-neutral workspace API that external agents authenticate against with single-use bootstrap tokens. simple-dispatcher is the runtime that executes the work. Its design is the unusual part: there is no in-process agent. A session is a detached terminal process driven entirely from the outside at hook boundaries, coordinating through a single JSON file as its IPC bus — which is what lets the same substrate run on a developer's Mac or a headless Linux box under tmux.

One sentence governs the whole thing: verified beats claimed. Work counts only when its evidence chain closes — CI green, merged, ledger discharged. The architecture is largely downstream of that: a review step that may only issue a conditional approval, a mandatory ledger of what CI did not prove, and a separate close-out step that owns the irreversible finish. "Done" is a claim; the artifact is the fact.

Progress

source git history, both repos
13/420/427/44/511/518/525/51/68/615/622/629/66/713/720/727/7
Tickets first appearing in a commit message, by week, across both repositories. Peak 163 in the week of 29 June; last four full weeks average 69.8/wk against 115.2/wk the prior four — -39%. The current partial week is excluded.

Delivery is down by about a third from its late-June peak, and the decline is steady rather than a single bad week. Two independent measures agree: tickets reaching code fell -39%, and mainline delivery units — one per merged pull request, immune to how commits are squashed — fell -37%.

What shipped is coherent. The runtime became more portable, more bounded and more observable: a tmux driver took the substrate off macOS onto headless Linux, autonomous runs gained an enforced task budget at the dispatch seam, per-task cost capture landed, and a Live Console timeline made in-flight execution legible. Alongside those, a broad run of incident fixes covered launch, wake, restart, transcript and completion recovery.

What did not change is as important. Work per ticket stayed flat — 1.18 to 1.22 mainline units per ticket, +4%. Tickets did not get bigger or harder. Fewer of them simply arrived at code.

One measurement gap is worth stating plainly. Of 1,112 tasks marked complete, 743 can be traced to a commit citing their identifier — leaving 369 that cannot. Some of that is legitimate: research notes, documentation, decisions, and work predating the current history. But the residue is unverifiable by the project's own standard, and the gap is not measured anywhere.

Why

researched · git, tickets, docs

The effort did not disappear. It changed what it was spent on, and the change is visible in what the code actually is.

1/68/615/622/629/66/713/720/727/7
Test and verification code as a share of all code written, by week, both repositories. It crosses 50% in the week of 6 July and has not gone back. In absolute terms production code more than halved — 26,072 → 12,086 lines/week — while test code held steady, within one per cent of where it started.

The same shift shows up in what the agents spend their sessions doing. Across 68 tickets and 500 worker sessions sampled from the live API, fewer than one session in four writes code:

23%Implementation117 of 500 sessions write code
54%Plan, review, close-out271 of 500 sessions · governance
12%Orchestration62 of 500 sessions · coordination
5.0 → 8.0Sessions per ticketmedian · week of 6 Jul → week of 27 Jul
44 → 63%Governance sharesame cohorts — the growth is all here
$26.01Median cost per taskmean $35.26 · $8.97–$89.08 · n=13

Four causes are visible in the record, and they are not equally weighted.

28–29 Jun

Two process changes landed 48 hours apart, at the exact peak

LIN-550 split close-out from review and introduced the "What CI Did Not Prove" ledger. The next day LIN-791 instructed the orchestrator to "decompose it into 3–6 ordered, self-contained beats" and "label every send beat N/M". The literal string beat N appears in zero commits before 1 July, then 9, then 18 a week. Both were deliberate quality decisions.

7 Jul

An external shock broke the dispatch mechanism

Claude Code's own prompt-injection defence began refusing Harbour's bootstrap-token handoff — the trust handshake the entire dispatch path depends on. Credential and broker work consumed 24% of all commits in the week of 13 July. This one is one-off and external, and roughly a fortnight of the window is attributable to it.

ongoing

The review loop does not converge

The project's own autopilot diagnosed this, unprompted, as a cross-task pattern — and four high-value tickets are parked at exactly that gate today, each carrying the same line: "No code has been changed at any point — the ticket is still at plan stage." These tickets consume sessions and never reach code, so they are invisible in every delivery measure above.

"the plan enumerates a class of call sites by hand-listing files; the reviewer verifies by running an independent behavior sweep and finds another member of the class the plan did not name. The revision adds the named ones, and the next sweep finds one more … the two-cycle bound is reached before the enumeration converges." — LIN-1408, autopilot comment, 2 August 2026
9 Jul

The instrument stopped

Harbour runs a periodical whose stated job is "a severity-ranked read of what has been dragging recent delivery". Its last run was 9 July. It reported "velocity hit a new all-time high … everything else is healthy and improving" — in the same week delivery began falling. Twelve lines earlier, the same document had written down the trap it then walked into:

"Git merge cadence tracks throughput, not forward delivery — so a rising commit count is never read as rising forward progress." — Recent Headwinds review, 9 July 2026 · the last one to run

Ten of fifteen periodicals currently read never. Nothing in the project's own record observes the decline.

The sharpest detail. The plan-review gate shipped with a recorded baseline and an explicit guard — "the step must not tax the throughput it exists to protect" — plus a follow-up ticket, LIN-1661, to re-read the number one cycle later. LIN-1661 is still Todo. The gate's own falsification test has never been run, and the window it would have covered is exactly the window delivery fell in.

The shape of the work

derived · 97% of all tasks

Harbour began in January 2026 as a read-only viewer for Linear projects, little more than a tree and a footer. Within weeks it started generating prompts, and the story becomes one of an app teaching itself to act: grounded prompts gave the work shape, a safe API gave it hands, and the autonomous run loop turned a person clicking dispatch into a fleet of agents dispatching each other. That autonomy exposed how much the system lied to itself, so the largest single investment became making the runner honest — watched by the observation surfaces and defended by the test suite. In parallel the foundations were generalised: the backend became swappable, then the agent engine, then the host, then the human. The remaining frontier is the north star itself — cost per verified task is only now becoming measurable, recurring self-review has stalled short of true cadence, and the steering layer is being rebuilt around planned voyages.

The product surface — 178/222 done. From a bare read-only Linear tree to a designed, themed, mobile-aware product. The product surface 178/222 · 80% Grounded prompts & routing — 97/132 done. Turning a ticket into a prompt an agent can act on without hallucinating. Grounded prompts & rou… 97/132 · 73% A safe API for agents — 83/118 done. The workspace API agents talk to, and the fight to stop credentials leaking through prompts. A safe API for agents 83/118 · 70% Beyond Linear — 82/114 done. Making the issue backend swappable behind a canonical model and capability flags. Beyond Linear 82/114 · 72% The autonomous run loop — 95/144 done. The layer that picks work, dispatches it, follows up, wakes sessions and escalates. The autonomous run loop 95/144 · 66% Making the runner honest — 93/139 done. The out-of-process state machine, and the campaign to stop it reporting done for unfinished work. Making the runner honest 93/139 · 67% Watching the swarm work — 59/87 done. Surfaces showing what agents are doing right now — sessions, telemetry, transcripts, live console. Watching the swarm work 59/87 · 68% More than one agent engine — 75/111 done. Breaking the assumption that 'the agent' means Claude Code. More than one agent engine 75/111 · 68% A test suite that gates — 75/130 done. Turning a flaky, mock-backed suite into something that can actually block a merge. A test suite that gates 75/130 · 58% Getting off the Mac — 37/76 done. Escaping AppleScript/iTerm so the dispatcher can run headless on Linux. Getting off the Mac 37/76 · 49% From one user to many — 52/83 done. Durable accounts and owner-scoped credentials — prerequisites for anyone but the author using this. From one user to many 52/83 · 63% Knowing where we're going — 78/114 done. Direction surfaces that decide what the autonomy points at. Knowing where we're g… 78/114 · 68% Recurring self-review — 64/94 done. Scheduled, grounded review tasks the system files against itself. Recurring self-review 64/94 · 68% Cost per verified task — 14/47 done. The north-star metric: what a verified task actually costs — and making it fall. Cost per verified task 14/47 · 30%
↓ depends on / came first bar delivered ÷ in scope ◐ active ○ emerging ✗ stalled
◐

The product surface

ACTIVE178/222 · 80%

Began in January as a project tree with features bolted on as hand-rolled HTML. From April the debt was paid deliberately: a shared page shell, a component library, a token/typography design system, a full UI audit and a page-level glow-up. What remains is polish and accessibility residue.

LIN-748 · LIN-953 · LIN-1042

◐

Grounded prompts & routing

ACTIVE97/132 · 73%

An early audit found the hand-written templates contradicted each other, and the AI meta-prompt path immediately drifted from them. The fix was structural: one shared grounding post-pass so both paths execute the same rules once — plus staleness re-grounding, the class check, plan-fidelity, and the review/close-out split.

LIN-435 · LIN-313 · LIN-550 · LIN-698

◐

A safe API for agents

ACTIVE83/118 · 70%

The secure proxy landed in March — the moment the viewer stopped being read-only. Security drove the arc from standing tokens in prompt text, to single-use bootstrap tokens, to a localhost broker where the agent never sees a bearer at all.

LIN-193 · LIN-376 · LIN-1375

◐

Beyond Linear

ACTIVE82/114 · 72%

Everything initially spoke Linear GraphQL directly. A canonical model came first, then the consumer API itself was routed through providers, retiring the raw passthrough. A writable local provider was pulled ahead of GitHub to unblock testing. Jira was filed in February and never started.

LIN-174 · LIN-306 · LIN-356 · LIN-1504

◐

The autonomous run loop

ACTIVE95/144 · 66%

It started as Foreman Mode — one session driving a project — retired in June for the autopilot proper: a dispatch queue with feedback, follow-ups as a first-class capability, then push-based wake rails. Most 2026 defects lived here. The frontier is now governance: an operator decision queue and bounded runs.

LIN-209 · LIN-826 · LIN-1721 · LIN-1751

◐

Making the runner honest

ACTIVE93/139 · 67%

The runner drives detached terminals coordinated only through a JSON file, so almost every failure was a lie about state: lost-write races, terminal markers firing early or never, stale waits wedging sessions forever. July's rebuild collapsed thirteen phases to eight. It remains the noisiest arc.

LIN-459 · LIN-549 · LIN-900 · LIN-1113

◐

Watching the swarm work

ACTIVE59/87 · 68%

The Pipeline page was the first attempt and was eventually deleted. The real answer was the Observation page, grouping dispatches into sessions, then a per-session page with a reply box so a human can answer an agent waiting on them. Remaining work is fidelity.

LIN-595 · LIN-1003 · LIN-1436

◐

More than one agent engine

ACTIVE75/111 · 68%

For most of the year the runner could only launch Claude Code. OpenCode arrived as a genuine second harness with its own per-session runner. The hard part was parity, not launching: OpenCode has no Stop hook, so heartbeats, sentinels and evidence mining were rebuilt. Cursor and Gemini are filed but unstarted.

LIN-393 · LIN-1077 · LIN-1404

◐

A test suite that gates

ACTIVE75/130 · 58%

Early specs asserted against mock branches inside production routes, so green meant very little. Every spec was migrated onto the real local provider, then a flakiness sweep, then a full-system hermetic suite that boots a real Harbour and drives the real dispatcher. CI still cannot be made a required check.

LIN-215 · LIN-798 · LIN-1490 · LIN-1580

◐

Getting off the Mac

ACTIVE37/76 · 49%

Execution was welded to macOS, with absurd failure modes — reapers closing terminals mid-conversation, terminology faults after a reboot. A driver port allowed Terminal.app, then kitty, finally tmux, which needs no display server. Cloud execution and a multi-box fleet are specified but entirely unstarted — the largest concentration of untouched work in the project.

LIN-767 · LIN-1782 · LIN-1301 · LIN-1792

◐

From one user to many

ACTIVE52/83 · 63%

The app originally equated a Linear user id with a human, which broke the moment a workspace could be GitHub-backed. Durable accounts came first, then preferences and tokens re-keyed onto them. The sharp edge was credentials: workers died whenever the owning human's session expired. The hosted path is barely begun.

LIN-1326 · LIN-1539 · LIN-1414

◐

Knowing where we're going

ACTIVE78/114 · 68%

It started as a deterministic roadmap with an LLM narrative on top, then a north-star layer and a radial view orienting tasks by compass bearing. The payoff came later: suggested next run, turning roadmap state into goal options a human can accept. The newest form is a human-ratified planning session that writes a multi-leg voyage.

LIN-222 · LIN-273 · LIN-603 · LIN-1811

✗

Recurring self-review

STALLED64/94 · 68%

A registry of periodicals whose dispatch generates a grounded task rather than a generic reminder, growing to eleven templates. June saw the peak — all eleven ran in one sweep. Since then the cadence has not recurred: two dated runs in July, eleven run-tickets cancelled. The trigger that would make it autonomous is still open.

LIN-315 · LIN-373 · LIN-1629

○

Cost per verified task

EMERGING14/47 · 30%

Cost was invisible for the first half of the year. Real capture arrived only in late July, when the runner began emitting per-session usage. The north-star derivation is in progress right now, blocked on a definition rather than on engineering. This is the youngest and thinnest arc, and almost everything that would make cost fall is still ahead.

LIN-418 · LIN-1425 · LIN-1625

Forecast

deterministic · arrival rate unmeasured

At the current rate of 69.8 tickets/week, the 549 open tasks represent 7.9 weeks of work — reaching 27 September 2026. At the prior rate it would be 4.8 weeks.

But that figure is a floor, and a receding one. Tickets are being created at about 190 a week against 69.8 reaching code — so the backlog grows by roughly 120 tasks every week at current rates. The burn-down above answers "how long if creation stopped", which it will not. And nothing in this workspace carries a due date — 0 of 1,810 tasks — so there is no schedule to be on or off. Every colour in this brief is flow health, never schedule health.

The more useful forward statement is about sequencing, not dates. Harbour's own north-star reading names the open question directly:

"Should active Product interface work continue, or be paused behind reliability incidents and the blocked cost-per-verified-task chain?" — Harbour roadmap digest, generated 2 August 2026, LLM-authored, marked fresh

That question has force because of a specific inversion. The project's stated headline metric is cost per verified task, visible and falling — and the task that would derive it from existing capture (LIN-1625) is in progress, stale, and blocking eight others. Meanwhile the raw data to compute it already exists and is queryable today: the figures in the metrics below were derived from the live API in about fifteen seconds.

Where the work sits

deterministic · task states
Harbour · 1,423 tasks in scope · 68% delivered
excludes 128 canceled & duplicate and 259 unassigned to a workstream — 1,423 + 259 + 128 = 1,810
├─◐Product398/507 · 79%
├─◐Autopilot, Recommendation & Prompt Engine136/218 · 62%
├─○Simple Dispatcher134/203 · 66%
├─◐Dispatch & Execution Runtime63/114 · 55%
├─○Quality, Periodicals & Measurement38/82 · 46%
├─✗DevOps & Tooling55/66 · 83%
├─○Platform Security, Robustness, Observability22/66 · 33%
├─◐Providers & API Unification51/64 · 80%
├─◐UX, Theme45/62 · 73%
├─✓Test Infrastructure: Local-Provider Migration28/28 · 100%
└─○The Ship's Biscuit4/13 · 31%

✓ complete   ◐ healthy   ○ at risk   ✗ stalled  — bar shows delivered ÷ in scope

Moving parts

deterministic · rule printed per card
✗

DevOps & Tooling

STALLED

11 open, 0 landings in 30d (last was 45d ago)

55 done11 open1 bugship 45d ago
○

Simple Dispatcher

AT RISK

5 open bugs (threshold 3); still shipping (17 landings/30d)

134 done69 open5 bugship 9d ago
○

Quality, Periodicals & Measurement

AT RISK

3 open bugs (threshold 3); still shipping (12 landings/30d)

38 done44 open3 bugship today
○

Platform Security, Robustness, Observability

AT RISK

open (44) exceeds delivered (22); still shipping (8 landings/30d)

22 done44 open2 bugship 2d ago
○

The Ship's Biscuit

AT RISK

open (9) exceeds delivered (4); still shipping (4 landings/30d)

4 done9 open0 bugship 22d ago
◐

Product

HEALTHY

73 landings in 30d, last 0d ago; 2 open bugs

398 done109 open2 bugship today
◐

Autopilot, Recommendation & Prompt Engine

HEALTHY

57 landings in 30d, last 1d ago; 1 open bugs

136 done82 open1 bugship 1d ago
◐

Dispatch & Execution Runtime

HEALTHY

17 landings in 30d, last 1d ago; 2 open bugs

63 done51 open2 bugship 1d ago
◐

Providers & API Unification

HEALTHY

4 landings in 30d, last 10d ago; 0 open bugs

51 done13 open0 bugship 10d ago
◐

UX, Theme

HEALTHY

38 landings in 30d, last 5d ago; 0 open bugs

45 done17 open0 bugship 5d ago
✓

Test Infrastructure: Local-Provider Migration

COMPLETE

0 open tasks; 28 delivered

28 done0 open0 bugship 40d ago

Risks & blockers

mixed · measured + judged
HIGH
Delivery is down about a third, and it is not a measurement artifact

Two independent measures agree: tickets reaching code, and mainline delivery units that are immune to merge policy. Work per ticket is flat, so this is fewer tickets converting — not bigger tickets.

tickets 115.2→69.8/wk (-39%) · mainline 135.5→85.2/wk (-37%) · per-ticket +4%

HIGH
The project cannot see its own slowdown

The periodical whose stated job is reading delivery drag last ran on 9 July, reporting an all-time high in the very week the decline began. Nothing in the project record observes what this brief measures.

last Recent Headwinds review 2026-07-09 · 24 days overdue · 10 of 15 periodicals read never

HIGH
The gate that may have caused it has never been tested

The plan-review gate shipped with a recorded baseline, an explicit guard that it must not tax the throughput it exists to protect, and a follow-up to re-read the number one cycle later. That follow-up has not run — and the unmeasured window is exactly the window delivery fell in.

LIN-1600 shipped 2026-07-26 · baseline follow-on ratio 0.2342 · LIN-1661 re-read still Todo

HIGH
The headline metric is blocked on a definition, not on engineering

Cost per verified task is the stated north star. The spend half is captured; the outcome half is not computable, because ledger discharge has no marker and no pull request is ever read. It waits on one human ruling.

LIN-1625 · in progress · unblocks 8 · 7 days stale · trial figures $17.83 vs $22.80 per task

MED
Tickets that consume sessions and never reach code

Four high-value tickets are parked at the plan-review bound, each carrying the line "No code has been changed at any point". They burn agent sessions and appear in no delivery measure — here or in Harbour.

LIN-1694 · LIN-1731 · LIN-1717 · LIN-1408 — 4+ sessions each, zero commits

MED
Coordination is now most of what the fleet does

Fewer than one agent session in four writes code; planning, review and orchestration take the rest. The wake traffic carrying that coordination is itself unowned.

implementation 117/500 sampled sessions = 23% · LIN-1749 reports wakes at 46% of all sessions

MED
Completed work that cannot be verified

By the project's own standard the artifact is the fact, yet a fifth of completed tasks have no commit citing them. Some are legitimately non-code; the residue is unmeasured.

743 of 1112 completed tasks traceable to a commit

MED
No schedule accountability is possible

Not a risk in the work but in the reporting. With no due dates anywhere, no stakeholder can be told whether anything is late, and no brief can honestly claim a project is on track.

0 of 1810 tasks carry a due date

Metrics

measured 3 August 2026
1,112Tasks deliveredof 1,810 · 61% · 149 canceled
69.8Tickets reaching codeper week · -39% vs prior 4 weeks
190Tickets createdper week · flat vs 193 prior
-54%Production code26,072 → 12,086 lines/wk · test -1%
549Open tasksacross 11 workstreams · 18 flagged bug
743Traceable to a commitof 1,112 completed · 67%

Cost figures are a sample of the 13 most recently landed tickets, each fully priced — no unpriced models, no missing telemetry. The endpoint returns a null total rather than a partial one, so these are complete or absent, never quietly understated.

How this was made

method · and one correction

This brief was assembled by an AI agent from live data, then attacked by another one. The attack succeeded, and the headline changed — so it is worth showing the work.

measureAll 1,810 tasks paged from the workspace API; full git history of both repositories; 500 agent sessions sampled for composition and cost.
explainSix independent agents on the slowdown: git forensics, a commit taxonomy, the theme graph, ticket bodies and comment threads, the project's own research docs, and one adversary.
refuteThe adversary's only job was to break the headline. It was given the claim, the data, and seven specific attacks to attempt.
correctIt succeeded on three of them. The original claim — "output barely fell, tickets just got heavier" — was withdrawn.

What the adversary found. The first draft claimed raw commits fell only 18% against a 49% fall in tickets, and concluded the work per ticket had inflated by ~77%. Three things were wrong with it. The two series were computed on different repository sets — a missing directory change had silently made one of them a duplicate of the other. The commit counts were inflated by a merge-policy shift: squashed pull requests contribute one commit, merged ones contribute a whole branch, and the mix moved from 92% squash to about 25%. And the four-week window was the only window length at which the two measures diverged at all — at three weeks and five weeks they track each other.

Re-measured on one consistent basis, using mainline units that are immune to merge policy, work per ticket is flat (+4%). Delivery genuinely fell by about a third. The composition shift and the session evidence — both independent of git — are what survived, and they are what this brief now rests on.

That is not an aside. Harbour's own thesis is that an optimizer scored on a surface it authored itself will drift, and that the fix is a check it cannot author. This page was written by an agent, checked by an agent it did not control, and corrected against its own first conclusion. The method is the subject.

What this brief cannot tell you

limits of the sources
  • Whether anything is late. No task or project carries a due date, so no schedule verdict is possible. Every status here is flow health, not schedule health.
  • What changed since the last brief. Report history stores narrative prose, not structured snapshots, and only the latest report is reachable — there is nothing to diff. "Progress" above means the trailing period, not the delta since last time.
  • Whether the review loop is worth its cost. This is the central open question and it is not answerable from here. The gates were added to stop defects escaping; measuring whether they did needs the follow-on-task ratio re-read that LIN-1661 owns and has never run. Reverts falling to near zero is suggestive, not proof.
  • Agent effort before 4 July. Session and cost telemetry expires at 30 days, so the session-composition evidence cannot be extended back into the June peak. The June-vs-July comparison exists for code, not for effort.
  • Total spend. Cost is per-task only, with no workspace aggregate. The dispatch history caps at 100 items with no pagination, and its unfiltered total of 7,129 does not reconcile with the per-status totals summing to 200 — so that number is deliberately not printed anywhere above.