Harbour Archive · Document 5 · An essay
The Cheap Ships
Five levers, what each one unlocks, and where they take Harbour. Told from the run as it stood on the evening of Friday 5 September 2026.
Part oneWhere the meter is
It is Friday evening, 5 September. The gauge on /kpis reads a little over seventy percent of the week's window spent. At the current burn the window runs dry on Sunday morning. It refills on Thursday. The fleet only moves while a tab is open, and each landed task costs about thirty dollars at API list rate.
That thirty dollars never reaches a bank account, because every session rides the subscription. The capacity review measured what that lane is worth: one prepaid pound buys thirty to sixty-five pounds of list-rate compute, when the week has headroom. Two of the four measured weeks did not. The cap binds, not the backlog.
John reframed the objective on 15 August, and it still holds. The goal is not to save money. It is verified outcomes per week under a fixed weekly budget. Tonight's addition is the second half of that sentence: run at full capacity, then step off the subscription altogether while keeping the books balanced.
The August reviews found the fact that makes the whole story possible. Ninety-five and a half percent of all spend is carried-context re-reads. An implementation leg carries about 190 times the context it was handed. The observer carried 1,420 times. The engines are not slow. The archive put it best: the engines run at the same speed they did in June, and what changed is wasted motion.
Then came the probes. Two real legs from the capacity day were re-run with the smallest context that still passed the same verifier. One cost twelve times less. The other cost forty-four times less. Removing a single ninety-three-token pointer to a file and a function more than doubled the cost of an otherwise identical handoff. On localized tasks Haiku matched Sonnet exactly from the same distillate, at a quarter of the price.
Every optimisation lined up in the wing attacks that one fact from a different side. Here they are in the order they compound.
Part twoFive levers
Lever one · uptimeThe observer that does not think
The first lever is already half-landed, and it is the one this week was spent on. The observer sweep classifies the whole fleet into seven lanes with no model call at all. The judgement pass spends one small stateless call only when the census changed, and its quiet path is deterministic, which the code describes as honest regardless of model behaviour. The Flight Companion's gate is a pure function that decides whether an auto-wake may spend anything.
The levers map priced this lever at the top of the table: coordination carry on a full-width day falls from roughly $580 to $30 or $60, because the orchestrator tier alone was almost a quarter of the day. What it unlocks is the thing John asked for first. The run keeps moving with the tab closed. He becomes the navigator rather than the presence required for every tick. The trust ladder he set, read-only first, then supervised writes, then unsupervised, is the shape the docket and its nine follow-ups now take. And it hands the next lever a clean problem: once coordination is nearly free, what remains is the worker legs, and worker spend is research.
Lever two · contextThe wiki
The design is a wiki base with deterministically generated stub files for everything that can be pulled out of code, plus meta-analysis indexes, so that a search surfaces almost everything a task needs and skips most of the research phase. The lineage is visible on John's GitHub: CodeWiki-Generator fingerprints a repository deterministically before any synthesis and already serves agents over MCP; Pith extracts deterministic facts into a node graph and exists to compress a codebase into prose an agent can use.
Harbour's own measurements say exactly why this works. The gap between finding an answer and emitting it was 2.2 times in cost and 2.8 times in turns. CLAUDE.md is 27.8 thousand tokens, is not in the cached preamble at all, enters mid-session in about a third of transcripts, and dropping it changed the verifier result not at all while cutting probe cost by a quarter. The outside evidence agrees. The AGENTS.md study found long context files add over twenty percent inference cost without raising success, while deterministic code maps such as CodeGraph self-report roughly half the tokens and tool calls.
A wiki regenerated from HEAD also makes the staleness check a diff rather than an expedition, and the cited-sweep rule becomes a query over an index. The long tail, nine sessions carrying thirty-one percent of a day, is the search tail. This shortens it. And it hands the next lever a task in the exact shape it eats: a pointer, a plan, a few files, a test command.
Lever three · priceDash
This is the lever that moves implementation off the frontier meter. Dash's pipeline is Understand, then Plan, then Implement as search-and-replace diffs applied through git in a worktree, then Verify against the test suite with up to three corrections. It decomposes automatically. It rejects tasks that exceed its complexity thresholds and returns guidance, which is the decomposition hint. Query mode runs the research phase alone and writes an answer, which is the read-only pass. The integration guide says the quiet part plainly: you are the approval layer, and Dash produces first versions. It defaults to a Gemini Flash-class model in the sub-dollar tier and takes any OpenRouter model or a local LM Studio endpoint.
Three of John's notes matter most here. Dash can call Dash, so an implementer can hand sub-slices to its cheaper self. Dash is built with Harbour, so the tool that cuts Harbour's costs is itself dispatched through Harbour's loop. And the discipline came from data: dash-analysis read four hundred of his own sessions and found error-recovery loops in three quarters of them and whitespace-broken string replacements in nearly a third. Diffs through git with a test loop are the direct answer.
Wiring it in is small. Simple-dispatcher's harness registry already holds the opencode template, with no Stop hook and process-exit liveness, and a Dash entry is one more row. Harbour already forwards harness, model and effort per kind, so the workspace setting is one field. LIN-1181, tracking Dash invocations and cost, has sat on the cheap-models roadmap since July. The old dash target that LIN-900 removed returns on the right axis this time, as a harness rather than a target.
Lever four · routingJudgement stays expensive on purpose
The mechanism is built: per-kind model, harness and now effort, resolved payload over workspace over default, with the effort flag landing on main today. The cheap-models roadmap says the remaining half is proof, with four acceptance tests: a cheap model finishes a real task unattended, you can see what it cost, the smart model catches the cheap one, and a run survives a hiccup.
The ceiling review supplies the routing rule. Work-shaped legs did not separate the tiers. Judgement-shaped legs did, and by depth of engagement rather than context need. So the frontier model buys plan-review, review, close-out and rulings. Cheap models buy implementation and lookups. Deterministic code buys the census, the grounding and the graph. The north star already forbids the alternative: pricing is policy, never judgement, and no agent reasons about price.
This week's news is on Harbour's side twice. Fable 5.1 cut cache reads to $0.25 per million, which is a direct discount on what Harbour spends the frontier on. And two papers plus Scale's standardized leaderboard show harness-induced variance exceeding model-induced variance, with the same Opus 4.5 scoring five points apart across three scaffolds and vendor tables running twenty points above harness-controlled runs. The harness is the product. That is Harbour's entire bet, now in print.
Lever five · runtimeHarbour OS
The last lever puts all of it on one floor. The landing page has the line: Harbour is the cockpit, Harbour OS is the workshop floor. Today Harbour spawns into it over an OSC escape as the local target. Browser-agent shows the philosophy the workstation grew from: inference in the browser over WebGPU, a deliberately tiny model, and the design rule that the model handles only small self-contained tasks while deterministic code manages the complexity. That is the same rule as the observer sweep, and the same rule as Dash.
When the wiki, Dash, the observer and the companion all run in Harbour OS, the macOS terminal-driver fragility that so much of simple-dispatcher exists to defend against simply stops being the runtime. Local inference finishes the picture. Apple Silicon on MLX now decodes a 35B mixture-of-experts coder at over a hundred tokens a second, and open-weight 30B models report 73 to 76 percent on SWE-bench Verified on a single 24GB card. John's earlier ruling that everything is paid by the user who starts it gains a new reading: the user's own GPU, at zero marginal cost.
Part threeWhat it adds up to
Here is the 100× claim decomposed against numbers Harbour has already measured or the market has already published.
| Factor | Grounding | Multiplier on cost per landed task |
|---|---|---|
| Coordination carry to the observer harness | Levers map: $580 of a $1,540 day falls to $30–60 | ≈1.5× on the whole day |
| Worker context, wiki plus pointer handoffs | Probes at 12× and 44.6×; search gap 2.2×; dropping CLAUDE.md 1.3× | 3–5× on worker legs |
| Model price tier, frontier to Flash-class or local | $10/$50 against $0.75/$3.75; DeepSeek V4 Flash at $0.14/$0.28; local at zero marginal | 10–30× on implementation legs |
| Rework and repetition, Dash's verify loop plus mutation checks | Buckets B3 and B4 were 15 percent of the day | ≈1.2× |
Multiplied, the implementation lane lands between roughly 50× and 250×. The judgement legs stay frontier-priced by design, so the blended figure on day one is more like 20× to 50×. It climbs toward 100× as decomposition pushes more of the work-product bucket, which was 56 percent of spend, into the cheap lane. So "comfortably 100×" is grounded for the lane that matters and a stretch-but-reachable number for the blend. Uptime is not in the table because it is a throughput multiplier rather than a cost one: a fleet that runs with the tab closed simply uses the whole week.
The exchange rate is where the story turns. The subscription is about £45 a week. Run to the cap, it buys on the order of $4,000 to $5,700 of list-rate compute across roughly 3.7 full-width days, and from 14 September the limits settle about seventeen percent below today's. At 10× the same week of work on the API costs a few hundred dollars, more than the plan but with no cap and no reset. At 100× it costs less than the plan fee.
Part fourThe tide, and how fast
The tide is running the same way, and faster than the plan needs.
| Signal | What it means for Harbour |
|---|---|
| The token expenditure index fell below $1 per million on 1 September, under half its summer peak | Dash's default tier gets cheaper without anyone touching Harbour |
| Epoch AI's estimate: price for fixed capability falls 9× to 900× a year by benchmark, about 40× a year at GPT-4 level | Harbour's 100× is its own work multiplied by a tide that adds roughly 10× a year on its own |
| Sub-dollar models now report 70 to 79 percent on SWE-bench Verified, where the frontier stood a year ago; the open-weight gap is three to five months | The Haiku-matches-Sonnet finding generalizes; the cheap lane is where last year's frontier now lives |
| METR's time-horizon doubling is 88.6 days; Mythos Preview's 50 percent horizon passed sixteen hours, past the point METR calls reliable | Longer autonomous legs mean fewer wakes per task, which the north star already counts as a tax |
| Subscription OAuth is blocked in third-party harnesses; Codex and Copilot moved to credits | The industry is pricing agents per token; BYOK through Dash is the grain, not against it |
How fast could it arrive? The observer docket and its follow-ups are in flight now. A Dash harness entry plus the workspace field is days of work once the hosted or BYOK path is stable, and LIN-1181 gives it a meter from the first run. A wiki regenerated from HEAD is a periodical, and CodeWiki-Generator already speaks MCP, so the lookup path exists. The routing evals are the cheap-models roadmap's four tests, run against the post-compaction base as the levers map always intended. Harbour OS as the runtime is the long pole, and it can absorb the others one at a time. Nothing in the chain waits on a model that does not exist yet.
Part fiveWhere it takes Harbour
Back to its own north star: completes verified backlog work at a cost and cadence a solo operator can sustain, funds itself doing it, and proves every word. At 100× the free tier becomes a line item the operator can afford to publish. A stranger trusting it in one sitting becomes a product rather than a hope. And because Dash is built with Harbour, the loop closes on itself: Harbour dispatches the work that makes Harbour cheaper, so each landed ticket lowers the price of the next.
The constraint then moves to the place the north star already named as scarce: operator minutes. The nineteen questions that parked overnight and turned out to be nearly all false alarms are the shape of the next bottleneck, which is why decision context and dependency impact was the right thing to prioritise first. When every ship is cheap, the pilot's attention is the only expensive thing in the harbour, and the surface that spends it well is the one that wins.
The folded-loop document has the closing line. The harbour is not where ships stop being ships. It is where they remain answerable to the land. Cheap ships, an expensive pilot, and the land still deciding what is worth doing and what counts as done. That is the game after it shifts.
Harbour Archive · document 5 · filed 2026-09-05 · served verbatim from docs/archive/5.html