Harbour v1: one person, one task, one proven merge
(the scope of the first version other people can use: drafted with agents from John's words, accepted only by the human, versioned, provenance recorded. The north star is the rule; this page is the scope. When they disagree, the north star wins and this page is revised.)
Why now
In John's words, 19 September 2026: "what I'm currently doing is building my ideal product, whereas now I want to make it so that other people can start using it; that's where Harbour really gets interesting." And: "this feels like the first time I can really see a clear idea of what this finished version 1 effectively is."
The papers of the same day say where to meet those people. Newcomers arrive at the ceiling the tools already offer, which has been "hand over one task" for a year; the gate they name is the cost of checking the work; adoption spreads by seeing a colleague's run; and provenance is the one thing shown to make checking cheaper. v1 is built on those four findings and on nothing that contradicts them.
Why this is the compelling v1
It is the smallest thing that keeps the whole promise. Harbour's first sentence is that it keeps human intent in command of AI execution and proves every word. One task, chosen by the person, run in the open, ending in a PR whose evidence they can check and a close-out only they can send, is that sentence with nothing removed. Everything Harbour does at rungs 5 and 6 for the operator is that same shape repeated. So v1 is not a cut-down Harbour; it is Harbour at the scale of one decision.
It lands where the evidence says people are, not where the model assumed. Three papers put the population's frontier at handing over one task and watching, put the stall one rung later at not watching, and found the rung the product had been positioned at, the saved prompt, measured by nobody. A v1 whose first delight is a prompt tailored to the person's own task, with the handover as the next rung on the same screen, meets both people without betting on either.
Almost all of it already exists. The loop that v1 sells ran three times on 19 September without a hand on the wheel: research, plan, implementation, review with sources re-fetched, close-out, merge. What is missing is the front door and the experience of one run seen by someone who did not build Harbour. That is a bounded list, and most of it already has a ticket.
It produces its own evidence and its own reach. The measures below include the one reading no published source has, which rung a person actually enters on. And a shared run is the strongest adoption signal the literature found: a colleague's run in front of you predicted a first agent run better than anything about you. v1 ships with its own instrument and its own way of spreading.
It is stable under what comes next. The tool ceiling has sat at rung 4 for a year. Nothing in v1 depends on the next model or the next vendor feature; the person, the task, the evidence and the click are the same whichever model runs the legs. That is why cheap models on the work-shaped legs is a hypothesis worth holding: the shape of v1 does not need the frontier to be good.
It changes what Harbour is for. Until now every improvement served one user, and the backlog reflected it: two thirds of it at rungs 5 and 6. A v1 that another person can use turns the same machinery into something that can be judged from outside, which is the only judgement that eventually matters.
It is honest about what it does not know. Whether people press Go or copy the prompt, what a run should cost, what the free number is, what the run should feel like: each is written below as a hypothesis with the way it gets tested. v1 is the first version that can be wrong in a way Harbour can see.
The bar, and what comes after it
docs/north-star.md: "Login → connected → first evidence-verified merge, without talking to the operator, is that step's bar."
v1 is built without depending on anyone else, and it is done at the dress rehearsal. In John's words: "we can build this v1 without a single other person and it will still be really, really good." The dress rehearsal is the first witness: "as soon as we reach the point where I can sign up a fresh account, go through all the steps and see a result, potentially with a repo that I've not used with Harbour before." The operator, as a stranger to his own product, on a repo it has never seen, all the way to the click. That is a measured artifact, and it is the gate.
Then the invited period. "Once I've done a sort of dress rehearsal, then it can start to be shared around, with me inviting other specific people to test it out." Specific people, measured by counts and by what they choose to share, during which the measures below start to mean something and the design track below gets its first outside eyes. Strangers reaching the bar unaided is the next step's bar, not this one's.
Who, and the job
A developer with a backlog in GitHub, Jira or Linear who wants one task done and proven. They may be hands-on today, saving the prompts that work, or they may arrive expecting to hand a task over. v1 offers the tailored next prompt first and the handover as the next rung beside it, so both find their step on the same screen.
The golden path
Eight things a person does, in order. What each step must achieve is stated; how it looks is the design track's to find.
- Log in. In John's words: "they arrive on harbour.cat, they log in with their email address."
- Connect. "They connect their GitHub repos and their Jira for work." A tracker and a repo, as one uninterrupted path.
- See the work. "They immediately see a lovely interface that clearly shows their work … or be shown the top of the stack for a task." The person recognises their own backlog and can see which task Harbour would take first.
- Start one task. "There's just a great big button that basically says Go." Go does what it says: one press starts the task, bounded to the PR, and the person can see that it has started. The reasoning for choosing the task, a prompt tailored to it, and a prominent copy sit right under Go. Beneath that sits the next rung of the same move with more handed over: run this step for me. Unlike Go, that is not the standing loop. Rungs not yet set up are shown, never hidden, and pressing one says what it needs. If a run is already live on the task, Go says so instead of offering a second start. Which way the person took the task is recorded, including a press on a rung not yet set up.
- Watch. "Just sees a lovely interface of the work in progress." At every moment the person knows what is happening, what comes next, what it has cost, and that nothing has been merged.
- Get the PR with its evidence. "Essentially just gets a PR at the end along with the ledger, all wonderfully displayed, and can easily understand the task, how it was run." The display exists to make checking cheap, not to make the run look good: what was proved, what was not, what the review read, each claim one step from its evidence.
- Close out. A user's run stops after review with the PR ready; close-out is theirs to send. It re-checks CI on the exact commit, settles the ledger, merges, and marks the task done. Harbour's server never calls GitHub's merge itself, and it refuses a close-out sent by the run itself. Beyond that, the stop rests on the run's instructions and on the runner's GitHub access: any session holding that access can still merge (credential-level enforcement is the LIN-3417 spike), and a run the owner launches without the stop closes itself out by design. Protecting the default branch on GitHub, for example by requiring an approving review, is what keeps an agent from merging on its own; Harbour advises it and does not check it.
- Share the link. "They can share these links like a CI/CD pipeline on the web." One run, one URL, readable by whoever holds it.
At step 4: the prompt comes first. The next prompt is AI-generated and tailored to the task, because templates are generic; it is free within the free tier's prompt allowance, for the person who wants to do it with their own agent. This is the rung-3 entry the chain under LIN-614 serves, and the first delight: log in, see the tasks, and within about three clicks a prompt is copied into an agent and a step of the task gets done. The handwritten templates stay available under "other prompts".
The run experience is its own work
This page does not specify the interface, and it does not price the design at nothing. In John's words: "there is work to be done to make an intuitive, elegant, wonderful interface for what the user is actually expecting and wanting; this may involve something akin to the passage planner but for the individual task, a chat, something like that; it's quite complex, but I can see there is a simple-ish solution in there."
What the experience must achieve, whatever shape it takes:
- The person is never asked something they cannot answer, and never asked twice. False-escalation rate is a headline measure.
- The person always knows the state of their task, what Harbour proposes next, and what has and has not happened to their repository.
- A person who has never seen Harbour reads a finished run and understands what was done, what was checked, and what they are trusting when they click.
- The hands-on person and the handover person are both at home on the same screen.
Candidate shapes, held as candidates: the passage planner's conversation at the scale of one task, where Harbour proposes the next step, the person says go or adjusts, and the evidence lands in the same thread; the expected steps shown as things to press or let run; a single page per run in the shape of a CI pipeline. The design track starts with the person's expectations, not with any of these, and it runs alongside the build from the first step of the order below, not after it. The first outside eyes on it are the invited testers.
The shape, as decided with John on 26 September 2026 (LIN-2947, after the paper docs/papers/harbour/what-a-run-must-show.md):
- A task has a task page; a run has a run page. The task page (
/workspace/:urlKey/task/:identifier, LIN-3329) is the page per task over its whole life: where it stands at a glance, the task's PR with its evidence and (for the owner) the merge click, every session it has had, its brief and recap, its description and comments, and its details, updating as it runs. The run page (/workspace/:urlKey/observation/session/:sessionId) follows one autopilot run and can span several tasks — it is the operator's session page translated out of Harbour's own vocabulary, in the shape of a CI pipeline page at its own URL, and its live view is that page unfinished; once finished it is a record. Both read in layers: a glance, a check, and a deep read where every claim is one click from its evidence. (The task page's PR section, merge click, description/comments and light polish are LIN-3340; its close-out box is keyed from the task's own tracker data, and the PR section and header carry it. The run page keeps its own session-keyed box and press client until LIN-3349 rebuilds it.) - While it runs: progress through the steps is the one number; money shows as information when it is known and is left out, never guessed, when it is not; active running time is kept apart from wall-clock time; there is no cap on the page, though the bound still guards the run. "Nothing merged" is said as a promise and, once there is a PR, as its live state.
- Chat is a first-class citizen, never the front page. It sits beside the page and explains; only the page acts, so anything agreed in chat reaches the run as a proposal on the page, recorded against its step.
- A run asks only for a decision only the person can make, or a blocker only they can clear, as one card pinned to the run page, with what happens if nobody answers. Harbour's own problems are reported, never asked. No email until runs are shown not to over-ask.
- A finished run leads with a paragraph written from its evidence, in the family of the brief and the recap, then asked, done and checked, then the ledger and the steps. Finished means finished: nothing straggles at the top, and anything that truly needs the person arrives as a question instead.
- The task page's share link is the same page seen as a guest, at an unguessable key in the URL (the mechanism of LIN-3073): a guest sees what the owner sees, with the controls hidden rather than disabled (John's ruling, 6 Oct), and links back into Harbour stay. The link can be revoked, and a revoked link shows Not Found. (Decided with the task page, LIN-3324; built in LIN-3330.)
- Observation is not a rung. It observes the runs at every rung, the way an operations platform grows from logs into a health view (ruling
lin2930-observation-rung: rung-agnostic).
The measure
- Time from login to the first Go.
- Share of people who reach the close-out click, and where the rest stop.
- Which mode each person used: took the prompt, took it a step at a time, or handed it over. This is the entry-rung measurement the papers could not find anywhere, taken on Harbour's own users.
- False-escalation rate on the run experience: every time it asks the person something, was the answer "why was I asked this?"
- Cost per verified task on the v1 lane, in the money actually spent.
In and out
In: the eight steps; the tailored prompt first, with the rungs beside it; the run experience as designed work; Harbour private by design, so the app never shows one account another account's work; a ring-fenced free tier; the hardening gate before anyone outside the operator is invited.
Out, as later steps rather than never: organisations and colleagues beyond the share link ("you add organisations, colleagues can share and watch the same thing" is the step after this one); machine and model configuration ("later they can specify things about the machines, the models"); passages and the standing loop for users, which stay the operator's; periodicals for users; Harbour OS. From the LIN-2947 sitting: a list of all a person's chats, and a chat pane docked beside the run; email for rulings; an admin view of aggregate numbers across the instance, and support access granted by a user; a merge button in Harbour with CI/CD integration, and a "close out when clean" rung.
A runner that serves people who are not the operator is out of v1's scope: where a run executes stays open by design (LIN-2938) — Harbour dispatches to whatever consumer polls the workspace, and v1 does not decide it. One rule stays: nobody else's work on a shared, uncapped key.
Hypotheses, stated as hypotheses
- Execution. Runs execute wherever a consumer polls the workspace (open by design, LIN-2938), with cheap models through OpenRouter on the work-shaped legs and the frontier only where judgement is bought. John: "even if right now you just plug in OpenRouter and everything runs through the cloud, that's still a really good product." On 2 Oct 2026 it ran once for real: in the
runner-witness-5db3ae2aworkspace on harbour.cat, the operator's own Claude Code session, given only a runner prompt copied from the site as its credential, took two queued items — one single step and one tiny Autopilot task that ran its own implementation → review → close-out workers, each item in its own subagent with brokered API access (witness LIN-3098e0b186df, verdict LIN-3098b450a067, close-outs LIN-309846a4d0bband LIN-3059e96347c4). That proved one witness, by the operator, on one tiny task, from a person's own machine — not a stranger, not a real repo task, and not cost at scale. The credential was owner-minted and expiring, withtake+dispatchgrants (LIN-3059, LIN-2884); an ordinaryreadWritetoken cannot take work, and how the task was taken is now recorded per account and task (LIN-2942, PR #1711). The first attempt, on 1 Oct, put a credential in prompt prose, fixed in LIN-3211; open residuals are LIN-3213, LIN-3230 and LIN-3236. - Pricing. The free tier is unlimited prompts plus a small, published number of cloud runs, per account per day. The paid tier is more runs on a small monthly subscription, or the person's own key. John: "potentially the cloud is the $5 subscription and you get copyable prompts, or something like that; we need to figure out the detail." The number is decided after the dress rehearsal, not before. The LIN-2947 tension — that the AI-tailored prompt counted against a daily prompt quota (
lib/free-tier-store.js) — is settled (LIN-3239): prompts are unlimited, the tailored prompt included; the bounded number is now runs per account per UTC day (FREE_TIER_RUN_LIMIT), and only a global hourly safety net can ever slow a prompt. - Entry. Most arrivals start with the tailored prompt and climb to the handover once they trust it. The draft of 19 September held the opposite; the first screen now bets this way, and the mode record, counting presses on rungs not yet set up, is how v1 finds out.
- The run experience. The shape above lets a person who did not build Harbour say what was done, what was checked and what they are trusting, from the paragraph and the page alone. The invited testers test that.
- Where a run executes, for a user. The person's own Claude Code session, opened with one copied prompt, takes their dispatched work and runs it through its own subagents, so every rung stays "copy this prompt". It worked once, locally; the spike LIN-3080 tests it against harbour.cat.
What exists and what is missing
As of 8a7f9c6e on main. Each row names where the gap already has a ticket.
| Step | Exists today | Gap for v1 |
|---|---|---|
| Log in with an email address | Linear OAuth, GitHub App installation ("Continue with GitHub"), Jira, personal access token (docs/architecture/auth.md) |
Email login does not exist in code, and the decision is already made: LIN-1892, ruled by John on 17 September 2026, makes the account the root and signs in by magic link or any linked provider identity. Build it. |
| Connect a tracker and a repo | Providers for Linear, GitHub, GitHub Projects, Jira and local; the GitHub App's repo picker | The connect flow as one uninterrupted path (LIN-614). |
| See the work | The stack (lib/task-stack.js), the home list and Swipe |
Home and Swipe stay. The opened task becomes one component shared by both: the next step with its reasoning and copy, the rungs, the proxy on and labelled plainly, the rest as a quiet row (LIN-2944). |
| Start one task | A scoped run exists: POST /api/proxy/autopilot/kickoff with an issue, a task budget, and a per-task worker-session budget (maxSessionsPerTask, LIN-2934) shown as "n of N" in the feed |
The three modes and the record of which was chosen; a stalled worker reported as stalled (LIN-2932); a failed press that says why (LIN-2933). LIN-2934 shipped the bound and the count for a scoped one-task run; v1 Go's own default value for maxSessionsPerTask is still undecided, deferred to LIN-2941. |
| A runner for people who are not the operator | Every run today executes on the operator's machine; the dispatcher already routes opencode through OpenRouter | Out of v1's scope, decided on 20 September 2026 (LIN-2938): where a run executes stays open by design, so Harbour dispatches to whatever consumer polls the workspace and v1 does not decide it; nothing in the epic depends on deciding it. The one rule that stays on the epic: nobody else's work on a shared, uncapped key. The runner milestone (LIN-2936) was cancelled on the same decision. |
| Watch the run | The Observation view and session pages, built for the operator | The task page follows one task over its whole life and the run page follows one autopilot run, both translated for a person who did not build Harbour, with the run's questions pinned to it (LIN-2948; task page LIN-3329). |
| The PR with its evidence | The review ledger and the close-out that consumes it | The display, each claim one step from its evidence; on the task page it is the Pull request section under the answer (LIN-3340), visible to guests, outside every repainted mount. |
| Close out | Close-out owns the merge today | For a user, the run stops after review and close-out is theirs to send (LIN-2949). On the task page the close-out box and press are keyed from the task's own tracker data (LIN-3340); the run page's session-keyed box awaits LIN-3349. |
| Share the link | A revocable token link to the task page: the owner mints it, a revoked link shows Not Found (LIN-3330, shipped); the old key-in-a-URL run share (LIN-3073) was removed by LIN-3325 | Sharing returns with the task page (LIN-3324); the guest sees the owner's page with the controls hidden. |
| Free tier | Runs per account per UTC day (FREE_TIER_RUN_LIMIT, lib/free-tier-store.js); prompts unlimited |
Show the run limit before Go; publish the number as a ring-fenced line item. |
| Paid tier | Nothing | Billing, after the dress rehearsal. |
| Hardening gate | The security review periodical; the north star's rule that secure defaults assume hostile input | Per-user token scoping, repo content treated as untrusted, Go's runs stop at the PR, with branch protection advised (the guards refuse the run's own close-out; the rest rests on prompts and GitHub access; credential-level enforcement is LIN-3417), private by design written down with a privacy page that tells the truth and a retention policy for every store (LIN-2954; the evidence-store ruling has since landed as lifetime retention, LIN-3157 B+D / LIN-3163), reviewed before the first invitation. |
| The measure | The funnel and the mode record (LIN-2952) | Login, connected, first Go and PR opened are derived from durable stores today (account.createdAt, account↔workspace membership edges, session-auth dispatches, the runner's [evidence] PR url); the merge-click step has no signal until LIN-2949's close-out calls the funnel-event store, so it reads "no signal available", never zero. One account's own five timestamps and mode are queryable at GET /workspace/:urlKey/api/milestone-funnel; the cross-account count is published on /kpis as milestoneFunnel. (LIN-1644 planned this line and was superseded.) |
Order
Risk first, and the design track alongside from the first step, because the experience is where the effort is and it cannot be left for last.
- Go as a bounded one-task run, with the three reporting fixes. The design track opens here.
- The run experience, the PR with its evidence, the close-out click, the share link.
- The funnel and the mode record.
- The free tier for runs.
- The hardening gate.
- The dress rehearsal. A fresh account, a repo Harbour has never seen, all eight steps, a result. v1 is done here.
- The invited period: specific people, measured by counts and by what they choose to share. Billing, then the paper on what they did.
Organisations, colleagues and configuration come after the invited period has taught what it can.
Provenance
- v1 draft, 19 September 2026: drafted by a Claude session from John Kershaw's words in the same planning conversation, under the authorship rule adopted on LIN-2926; the sentences of his behind each clause are quoted in place. The rung ordering (Go first, the prompt under it) and the risk-first order were the session's proposals, which John accepted in conversation the same day.
- Second draft, the same day, from John's review of the first: the bar moved from strangers reaching the milestone to the dress rehearsal, with invited testers as the period after ("we can build this v1 without a single other person and it will still be really, really good"); the interface un-specified and the run experience named as its own track of work; the case for why this is the compelling v1 added.
- Accepted by John Kershaw on 19 September 2026 by merging PR #1516. Two rows of the gaps table corrected the same evening from the tracker, after an adversarial review of the filing plan: the email door is already ruled (LIN-1892), and the runner is an open decision between the Linux box (LIN-1781/1785) and cloud execution (LIN-1301), with the credential lane the question either path must answer.
- Corrected 26 September 2026, per LIN-2938 (recorded from the conversation with John on 20 September 2026): the runner for people who are not the operator left v1's order and scope. Where a run executes stays open by design — Harbour dispatches to whatever consumer polls the workspace, and v1 does not decide it — and the one rule that stays on the epic is that nobody else's work runs on a shared, uncapped key. The execution hypothesis was corrected on the same basis: it had stated a cloud runner Harbour operates, and now says runs execute wherever a consumer polls the workspace.
- Revised 26 September 2026 from the LIN-2947 sitting between John and a Claude session, after the paper
docs/papers/harbour/what-a-run-must-show.md: the run experience's shape recorded under "The run experience is its own work"; the tailored prompt made the first delight, with the handover as the next rung (steps 4 and the aside under it, the entry hypothesis reversed and marked as such); step 7 changed from clicking merge to sending close-out, John's reading being that close-out is Harbour's version of the final human gate; the invited period's "watched" replaced by counts and what testers choose to share, under Harbour private by design; the later steps from the same sitting added to Out. - Recorded 2 October 2026, from the hosted runner witness in
runner-witness-5db3ae2aon harbour.cat (LIN-3237, the Go leg's last making-port item under LIN-3099): the Execution hypothesis now states what that one run proved — one witness, by the operator, on one tiny task, from a person's own machine — and what it did not — a stranger, a real repo task, or cost at scale. Evidence: witness LIN-3098e0b186df, verdict LIN-3098b450a067, close-outs LIN-309846a4d0bband LIN-3059e96347c4.