Harbour is generating work as fast as ever and converting a third less of it into shipped code. That gap is the story.
Tickets reaching code fell 39% (115.2→69.8/wk) and mainline delivery units fell 37%, while ticket creation held flat at about 190/wk — so conversion, not capacity, is what changed. Work per ticket is flat (+4%). Production code fell 54% while test code held steady. 1 workstream stalled, 4 under pressure, 5 healthy. No schedule risk is measurable: 0 of 1,810 tasks carry a due date.
Harbour is a provider-agnostic control plane for AI-augmented software development. It reads an issue backend, generates a grounded prompt for a specific task, dispatches that prompt to an AI agent, and verifies the result against real evidence. It began as a read-only Linear project tree and grew into the cockpit for that loop.
The system is two repositories. Harbour is the web control plane — provider adapters for Linear, GitHub Issues, GitHub Projects and a local store; a dual-path prompt engine; a dispatch queue; and a source-neutral workspace API that external agents authenticate against with single-use bootstrap tokens. simple-dispatcher is the runtime that executes the work. Its design is the unusual part: there is no in-process agent. A session is a detached terminal process driven entirely from the outside at hook boundaries, coordinating through a single JSON file as its IPC bus — which is what lets the same substrate run on a developer's Mac or a headless Linux box under tmux.
One sentence governs the whole thing: verified beats claimed. Work counts only when its evidence chain closes — CI green, merged, ledger discharged. The architecture is largely downstream of that: a review step that may only issue a conditional approval, a mandatory ledger of what CI did not prove, and a separate close-out step that owns the irreversible finish. "Done" is a claim; the artifact is the fact.
Delivery is down by about a third from its late-June peak, and the decline is steady rather than a single bad week. Two independent measures agree: tickets reaching code fell -39%, and mainline delivery units — one per merged pull request, immune to how commits are squashed — fell -37%.
What shipped is coherent. The runtime became more portable, more bounded and more observable: a tmux driver took the substrate off macOS onto headless Linux, autonomous runs gained an enforced task budget at the dispatch seam, per-task cost capture landed, and a Live Console timeline made in-flight execution legible. Alongside those, a broad run of incident fixes covered launch, wake, restart, transcript and completion recovery.
What did not change is as important. Work per ticket stayed flat — 1.18 to 1.22 mainline units per ticket, +4%. Tickets did not get bigger or harder. Fewer of them simply arrived at code.
One measurement gap is worth stating plainly. Of 1,112 tasks marked complete, 743 can be traced to a commit citing their identifier — leaving 369 that cannot. Some of that is legitimate: research notes, documentation, decisions, and work predating the current history. But the residue is unverifiable by the project's own standard, and the gap is not measured anywhere.
The effort did not disappear. It changed what it was spent on, and the change is visible in what the code actually is.
The same shift shows up in what the agents spend their sessions doing. Across 68 tickets and 500 worker sessions sampled from the live API, fewer than one session in four writes code:
Four causes are visible in the record, and they are not equally weighted.
LIN-550 split close-out from review and introduced the "What CI Did Not Prove"
ledger. The next day LIN-791 instructed the orchestrator to "decompose it into
3–6 ordered, self-contained beats" and "label every send beat N/M".
The literal string beat N appears in zero commits before 1 July,
then 9, then 18 a week. Both were deliberate quality decisions.
Claude Code's own prompt-injection defence began refusing Harbour's bootstrap-token handoff — the trust handshake the entire dispatch path depends on. Credential and broker work consumed 24% of all commits in the week of 13 July. This one is one-off and external, and roughly a fortnight of the window is attributable to it.
The project's own autopilot diagnosed this, unprompted, as a cross-task pattern — and four high-value tickets are parked at exactly that gate today, each carrying the same line: "No code has been changed at any point — the ticket is still at plan stage." These tickets consume sessions and never reach code, so they are invisible in every delivery measure above.
Harbour runs a periodical whose stated job is "a severity-ranked read of what has been dragging recent delivery". Its last run was 9 July. It reported "velocity hit a new all-time high … everything else is healthy and improving" — in the same week delivery began falling. Twelve lines earlier, the same document had written down the trap it then walked into:
Ten of fifteen periodicals currently read never. Nothing in the
project's own record observes the decline.
The sharpest detail. The plan-review gate shipped with a recorded baseline and an explicit guard — "the step must not tax the throughput it exists to protect" — plus a follow-up ticket, LIN-1661, to re-read the number one cycle later. LIN-1661 is still Todo. The gate's own falsification test has never been run, and the window it would have covered is exactly the window delivery fell in.
Harbour began in January 2026 as a read-only viewer for Linear projects, little more than a tree and a footer. Within weeks it started generating prompts, and the story becomes one of an app teaching itself to act: grounded prompts gave the work shape, a safe API gave it hands, and the autonomous run loop turned a person clicking dispatch into a fleet of agents dispatching each other. That autonomy exposed how much the system lied to itself, so the largest single investment became making the runner honest — watched by the observation surfaces and defended by the test suite. In parallel the foundations were generalised: the backend became swappable, then the agent engine, then the host, then the human. The remaining frontier is the north star itself — cost per verified task is only now becoming measurable, recurring self-review has stalled short of true cadence, and the steering layer is being rebuilt around planned voyages.
Began in January as a project tree with features bolted on as hand-rolled HTML. From April the debt was paid deliberately: a shared page shell, a component library, a token/typography design system, a full UI audit and a page-level glow-up. What remains is polish and accessibility residue.
LIN-748 · LIN-953 · LIN-1042
An early audit found the hand-written templates contradicted each other, and the AI meta-prompt path immediately drifted from them. The fix was structural: one shared grounding post-pass so both paths execute the same rules once — plus staleness re-grounding, the class check, plan-fidelity, and the review/close-out split.
LIN-435 · LIN-313 · LIN-550 · LIN-698
The secure proxy landed in March — the moment the viewer stopped being read-only. Security drove the arc from standing tokens in prompt text, to single-use bootstrap tokens, to a localhost broker where the agent never sees a bearer at all.
LIN-193 · LIN-376 · LIN-1375
Everything initially spoke Linear GraphQL directly. A canonical model came first, then the consumer API itself was routed through providers, retiring the raw passthrough. A writable local provider was pulled ahead of GitHub to unblock testing. Jira was filed in February and never started.
LIN-174 · LIN-306 · LIN-356 · LIN-1504
It started as Foreman Mode — one session driving a project — retired in June for the autopilot proper: a dispatch queue with feedback, follow-ups as a first-class capability, then push-based wake rails. Most 2026 defects lived here. The frontier is now governance: an operator decision queue and bounded runs.
LIN-209 · LIN-826 · LIN-1721 · LIN-1751
The runner drives detached terminals coordinated only through a JSON file, so almost every failure was a lie about state: lost-write races, terminal markers firing early or never, stale waits wedging sessions forever. July's rebuild collapsed thirteen phases to eight. It remains the noisiest arc.
LIN-459 · LIN-549 · LIN-900 · LIN-1113
The Pipeline page was the first attempt and was eventually deleted. The real answer was the Observation page, grouping dispatches into sessions, then a per-session page with a reply box so a human can answer an agent waiting on them. Remaining work is fidelity.
LIN-595 · LIN-1003 · LIN-1436
For most of the year the runner could only launch Claude Code. OpenCode arrived as a genuine second harness with its own per-session runner. The hard part was parity, not launching: OpenCode has no Stop hook, so heartbeats, sentinels and evidence mining were rebuilt. Cursor and Gemini are filed but unstarted.
LIN-393 · LIN-1077 · LIN-1404
Early specs asserted against mock branches inside production routes, so green meant very little. Every spec was migrated onto the real local provider, then a flakiness sweep, then a full-system hermetic suite that boots a real Harbour and drives the real dispatcher. CI still cannot be made a required check.
LIN-215 · LIN-798 · LIN-1490 · LIN-1580
Execution was welded to macOS, with absurd failure modes — reapers closing terminals mid-conversation, terminology faults after a reboot. A driver port allowed Terminal.app, then kitty, finally tmux, which needs no display server. Cloud execution and a multi-box fleet are specified but entirely unstarted — the largest concentration of untouched work in the project.
LIN-767 · LIN-1782 · LIN-1301 · LIN-1792
The app originally equated a Linear user id with a human, which broke the moment a workspace could be GitHub-backed. Durable accounts came first, then preferences and tokens re-keyed onto them. The sharp edge was credentials: workers died whenever the owning human's session expired. The hosted path is barely begun.
LIN-1326 · LIN-1539 · LIN-1414
It started as a deterministic roadmap with an LLM narrative on top, then a north-star layer and a radial view orienting tasks by compass bearing. The payoff came later: suggested next run, turning roadmap state into goal options a human can accept. The newest form is a human-ratified planning session that writes a multi-leg voyage.
LIN-222 · LIN-273 · LIN-603 · LIN-1811
A registry of periodicals whose dispatch generates a grounded task rather than a generic reminder, growing to eleven templates. June saw the peak — all eleven ran in one sweep. Since then the cadence has not recurred: two dated runs in July, eleven run-tickets cancelled. The trigger that would make it autonomous is still open.
LIN-315 · LIN-373 · LIN-1629
Cost was invisible for the first half of the year. Real capture arrived only in late July, when the runner began emitting per-session usage. The north-star derivation is in progress right now, blocked on a definition rather than on engineering. This is the youngest and thinnest arc, and almost everything that would make cost fall is still ahead.
LIN-418 · LIN-1425 · LIN-1625
At the current rate of 69.8 tickets/week, the 549 open tasks represent 7.9 weeks of work — reaching 27 September 2026. At the prior rate it would be 4.8 weeks.
But that figure is a floor, and a receding one. Tickets are being created at about 190 a week against 69.8 reaching code — so the backlog grows by roughly 120 tasks every week at current rates. The burn-down above answers "how long if creation stopped", which it will not. And nothing in this workspace carries a due date — 0 of 1,810 tasks — so there is no schedule to be on or off. Every colour in this brief is flow health, never schedule health.
The more useful forward statement is about sequencing, not dates. Harbour's own north-star reading names the open question directly:
That question has force because of a specific inversion. The project's stated headline metric is cost per verified task, visible and falling — and the task that would derive it from existing capture (LIN-1625) is in progress, stale, and blocking eight others. Meanwhile the raw data to compute it already exists and is queryable today: the figures in the metrics below were derived from the live API in about fifteen seconds.
✓ complete ◐ healthy ○ at risk ✗ stalled — bar shows delivered ÷ in scope
11 open, 0 landings in 30d (last was 45d ago)
5 open bugs (threshold 3); still shipping (17 landings/30d)
3 open bugs (threshold 3); still shipping (12 landings/30d)
open (44) exceeds delivered (22); still shipping (8 landings/30d)
open (9) exceeds delivered (4); still shipping (4 landings/30d)
73 landings in 30d, last 0d ago; 2 open bugs
57 landings in 30d, last 1d ago; 1 open bugs
17 landings in 30d, last 1d ago; 2 open bugs
4 landings in 30d, last 10d ago; 0 open bugs
38 landings in 30d, last 5d ago; 0 open bugs
0 open tasks; 28 delivered
Two independent measures agree: tickets reaching code, and mainline delivery units that are immune to merge policy. Work per ticket is flat, so this is fewer tickets converting — not bigger tickets.
tickets 115.2→69.8/wk (-39%) · mainline 135.5→85.2/wk (-37%) · per-ticket +4%
The periodical whose stated job is reading delivery drag last ran on 9 July, reporting an all-time high in the very week the decline began. Nothing in the project record observes what this brief measures.
last Recent Headwinds review 2026-07-09 · 24 days overdue · 10 of 15 periodicals read never
The plan-review gate shipped with a recorded baseline, an explicit guard that it must not tax the throughput it exists to protect, and a follow-up to re-read the number one cycle later. That follow-up has not run — and the unmeasured window is exactly the window delivery fell in.
LIN-1600 shipped 2026-07-26 · baseline follow-on ratio 0.2342 · LIN-1661 re-read still Todo
Cost per verified task is the stated north star. The spend half is captured; the outcome half is not computable, because ledger discharge has no marker and no pull request is ever read. It waits on one human ruling.
LIN-1625 · in progress · unblocks 8 · 7 days stale · trial figures $17.83 vs $22.80 per task
Four high-value tickets are parked at the plan-review bound, each carrying the line "No code has been changed at any point". They burn agent sessions and appear in no delivery measure — here or in Harbour.
LIN-1694 · LIN-1731 · LIN-1717 · LIN-1408 — 4+ sessions each, zero commits
Fewer than one agent session in four writes code; planning, review and orchestration take the rest. The wake traffic carrying that coordination is itself unowned.
implementation 117/500 sampled sessions = 23% · LIN-1749 reports wakes at 46% of all sessions
By the project's own standard the artifact is the fact, yet a fifth of completed tasks have no commit citing them. Some are legitimately non-code; the residue is unmeasured.
743 of 1112 completed tasks traceable to a commit
Not a risk in the work but in the reporting. With no due dates anywhere, no stakeholder can be told whether anything is late, and no brief can honestly claim a project is on track.
0 of 1810 tasks carry a due date
Cost figures are a sample of the 13 most recently landed tickets, each fully priced — no unpriced models, no missing telemetry. The endpoint returns a null total rather than a partial one, so these are complete or absent, never quietly understated.
This brief was assembled by an AI agent from live data, then attacked by another one. The attack succeeded, and the headline changed — so it is worth showing the work.
What the adversary found. The first draft claimed raw commits fell only 18% against a 49% fall in tickets, and concluded the work per ticket had inflated by ~77%. Three things were wrong with it. The two series were computed on different repository sets — a missing directory change had silently made one of them a duplicate of the other. The commit counts were inflated by a merge-policy shift: squashed pull requests contribute one commit, merged ones contribute a whole branch, and the mix moved from 92% squash to about 25%. And the four-week window was the only window length at which the two measures diverged at all — at three weeks and five weeks they track each other.
Re-measured on one consistent basis, using mainline units that are immune to merge policy, work per ticket is flat (+4%). Delivery genuinely fell by about a third. The composition shift and the session evidence — both independent of git — are what survived, and they are what this brief now rests on.
That is not an aside. Harbour's own thesis is that an optimizer scored on a surface it authored itself will drift, and that the fix is a check it cannot author. This page was written by an agent, checked by an agent it did not control, and corrected against its own first conclusion. The method is the subject.
100 items with no pagination, and its unfiltered
total of 7,129 does not reconcile with the per-status totals summing to
200 — so that number is deliberately not printed anywhere above.