# The ladder: how people come to trust AI agents with their work

*Descriptive, not normative. This document models where developers are and what Harbour owes them at each step. Agents may maintain it and revise it from evidence; the first paper on it is LIN-2925. The rule it implies lives in the north star, which only a human accepts.*

Drafted 19 September 2026 by a Claude session in conversation with John Kershaw, from his description of the ladder. His words are quoted where the rungs are defined. Revised the same day from the first paper (`docs/papers/harbour/developer-adoption-ladder.md@7f1da76f`), with the amendment accepted by John, and again the same day from the second and third papers (`docs/papers/harbour/developer-adoption-ladder-earlier-dates.md@433f629e`, `docs/papers/harbour/what-lowers-the-verification-cost.md@433f629e`) under the document's own rule.

## Why a ladder

Harbour exists to keep human intent in command of AI execution, so it has to model the human. People do not arrive at "feed tasks in and they land" in one step. They climb, and each step is gated by **trust**: have I seen it be right enough times to let it do the next step without watching? In John's words: *"it's more a single product, but we have to factor in how people will actually use it and how far along they'll go, and that's based on trust and budget and other factors like that."*

The ladder is one person's path, and the population is not one person. In John's words, after the first paper: *"the rungs are my personal journey, which seems to be happening to everyone, though of course as people start these journeys at different points in time the models' behaviour is actually different."* Each cohort starts on a higher rung because the tools moved between their start dates, and later entrants arrive beside a colleague rather than by climbing. So "where developers are" is a distribution over cohorts, not a queue, and the rungs themselves are defined against the tools of the day.

The tools set the ceiling, and the ceiling moves in steps. The highest rung any generally available tool offered was 2 at September 2024, became 4 on 25 September 2025 when GitHub's Copilot coding agent went GA, and had not moved by March 2026; agent teams shipped in February 2026 behind an environment variable. Over that six-month stall the population kept moving up into rung 4. So a cohort's entry rung is the ceiling on its start date, and most of the movement between ceilings is people filling rungs that already exist, not new rungs appearing.

What the published sources name as the gate is narrower than "trust and budget". Every source that names a cause names the cost of verifying the work; budget barely appears at the level of the individual developer. Budget is the operator's gate and it is real at rung 6. Below that, the price a person pays is the price of checking.

## The rungs

| Rung | The person | Harbour gives | Evidence at this rung | Gate to the next rung |
|---|---|---|---|---|
| 1 | Asks a question, copies the answer, carries on | Nothing yet | Their own eyes | Trust in one answer |
| 2 | Opens Claude Code and walks through the work by hand, meandering | Nothing yet | Their own eyes | Learns the sharp edges; learns to trust |
| 3 | Bundles repeatable steps into prompts and saves them | The grounded next prompt for a task | The tracker: the issue moved after the prompt was copied | "Run the next one for me?" |
| 4 | Lets Harbour run one task, and watches | Dispatch, a PR, CI, the review ledger, one approval click before merge | The PR and the ledger | Checking a landing costs less than doing it |
| 5 | Ratifies a passage of several tasks | Legs, task budgets, a landing report, rulings | The landing report | Rulings at a rate they can sustain |
| 6 | Feeds tasks in | The always-on loop, a forecast, KPIs | Cost per verified task; the forecast scored against actuals | None; this is the top |

John's description of rungs 1 to 3, 19 September: *"First people will ask Claude a question, copy the answer, carry on with the day. Next people open Claude Code and essentially walk through their work, meandering left and right, and eventually arrive at a place, and this is where they learn the sharp edges of Claude, this is where they learn to trust Claude. Then people start bundling up tasks: if they know that they need to go through four or five steps and it's a repeatable process they'll start making prompts, they get really good results, they save their prompts in the hope that they can use them later. This is where the majority of developers today are."*

And of what Harbour does with that: *"Harbour's first parlour trick is essentially giving them those prompts. They log in, they get their prompt, off they go. That is the core of Harbour, and everything else nearby exists to accelerate that: that's the autopilot, that's the automatic dispatch system. All of those exist to remove the friction of copying and pasting and then triggering the next actions. Harbour then extends this beyond where most developers are to the next level, which is multiple tasks as part of a passage, and leads eventually to an enormous, constantly running Harbour where you simply feed in tasks and they are done: complete, quick, no fuss, all correct, and land."*

## Where the population sits, on published evidence

From the first paper, each figure with its source's date. Rung 1 and above, 84 to 90% (mid-2025 to early 2026). Rung 2 and above, 62% (late 2025) rising to about 90% (mid-2026, on a broad definition of "agent"). Rung 3, unmeasured: no dated survey with a denominator asks whether developers save their prompts, so the "majority of developers today" line above is John's reading and stays labelled as such until a paper or a Harbour measurement says otherwise. Rung 4 and above, 31% (mid-2025) to 59% (April 2026), on the same Stack Overflow question asked twice. Rung 5 and above, about a fifth of all developers ever let an agent run unattended (April 2026). Rung 6, no figure; the defensible reading is low single digits or less.

The one published within-user climb is Anthropic's auto-approve share: about 20% of sessions for a new Claude Code user, over 40% at 750 sessions. The same series shows experienced users interrupting more often, because they stop approving each action up front and step in when something goes wrong.

At earlier dates the picture has the same shape, lower down. Rung 1 and above ran from roughly 60 to 85% at September 2024 depending on the instrument, to 84 to 90% a year later. Rung 4 had no population figure at all at September 2024. Rungs 3, 5 and 6 were unmeasured at every date; no survey has ever asked a developer population whether it saves a prompt. And one question at one date, "uses AI", spans 42% to 97% across instruments, so a difference between two dates smaller than that is not evidence of movement on its own.

## What follows from the ladder

**The instruments are the handrail.** Verified-beats-claimed, the ledger, rulings and cost per verified task are what make each climb safe to attempt. Each belongs to one rung. A person sees their rung plus one: a rung-3 developer needs "was that prompt good", not cost per verified task. Of the handrails, the one with a measured grip is provenance: showing where a claim came from cut verification time in the only controlled study that measured it. No published study shows that a stronger verification signal, green CI or a merged PR, changes how much a developer supervises; two that tried found behaviour barely moved. Verified-beats-claimed is a rule about what counts as done. It is not evidence that a verified artifact buys a person's trust, and this document does not assume it does.

**Rung-3 evidence is free.** Harbour reads the tracker, so it can see whether an issue moved after a prompt was copied without asking the user anything. That is the rung-3 form of external evidence over self-report.

**The big jump is later than it looks.** v1 put it at 3 to 4. The population does not stall there: agent use nearly doubled in eleven months, and a large organisation crossed that step in a quarter. Where the numbers fall off a cliff is between using an agent and not watching it: 59% use one, about a fifth ever let it run unattended, and standing loops are unmeasured. Rung 4 still has to be tiny, one task, one bounded run, the merge waiting for their click, the transcript in view, but the reason has changed. People do not refuse to reach rung 4; what they refuse is to stop watching, and the tiny shape is what lets them watch cheaply. The instruments that make watching cheap are the 4-to-5 handrail. Harbour already built this shape for the Flight Companion (LIN-2627: read-only, then supervised writes, then unattended). The same three steps are a user's on-ramp.

**Adoption is social, and lower-rung fluency does not carry.** The strongest disconfirming source in the first paper, a rollout across tens of thousands of Microsoft engineers, found that trying an agent was predicted by whether nearby colleagues already had, that tenure barely mattered, and that prior IDE-assistant use raised trying and lowered retention. If that generalises, the rung-4 on-ramp is partly a distribution problem: a person's first agent run is more likely to happen beside someone who already did one than at the end of a private climb. What that does to a product that meets people at rung 3 is John's call; LIN-2930 carries the question.

**Budget lines up with rungs.** Rung 3 costs almost nothing to serve, because the handwritten prompt templates are deterministic. The AI recommendation is the first BYOK or free-tier step. Rung 4 is the first time a task costs money on someone's key. Budget is what the operator feels at rung 6; the sources do not name it as a gate below that.

**The runner is the rung-4 on-ramp for people who are not the operator.** A rung-4 user will not install simple-dispatcher on their own machine. Machines as account objects (LIN-2883) and cloud execution (LIN-1301) exist for that step, and belong after rung-3 entry and before anyone is asked to climb.

**Rung 6 has one user today; rung 5 has more company than v1 assumed.** The operator uses rungs 5 and 6 to build Harbour; that is dogfooding and it is valid. About a fifth of developers sometimes let an agent run unattended, so a passage is a step people already take, not an exotic one. On 19 September 2026 roughly two-thirds of the open backlog was rung 5 and 6 work (the dispatcher-substrate, prompt-engine, cost-economy, proxy-api, rulings, operating-model, flight-companion and periodicals fronts). It is rationed to what raises verified tasks per week, not to new capability.

## What to measure

A funnel: people per rung, time on rung, and why they stop. LIN-1644 (time-to-trust instrumentation) is the ticket. The rung-3 number is the tracker-movement form, the issue moved after the prompt was copied, not prompts copied: an activity count can read green while the outcome is wrong, and METR's trial is the evidence, with developers who reported a 20% speedup measured 19% slower. For rungs 4 to 6 Harbour already has externally witnessed signals and should prefer them to any self-reported rung: the merged PR and discharged ledger at rung 4, rulings per landing report at rung 5, cost per verified task and forecast against actuals at rung 6. For the climb itself, the share of runs in each permission mode by session count would be Harbour's analogue of Anthropic's auto-approve curve. It is not computable today: permission mode is a single hardcoded launch flag in the dispatcher, stored nowhere, and supervision time is not recorded. The third paper's proposals line asks for both fields before the measurement is attempted. v2 of this document said the data already existed, and was wrong.

## Where the estimate of "where developers are" comes from

v1, from John's own reading of the developers he works with. v2, from the first paper, `docs/papers/harbour/developer-adoption-ladder.md` (LIN-2925), which mapped fifteen published sources onto these rungs and named the strongest disconfirming source. v3, from the second and third papers: `docs/papers/harbour/developer-adoption-ladder-earlier-dates.md` (LIN-2929), the same question at three earlier dates with the highest rung any tool offered at each, and `docs/papers/harbour/what-lowers-the-verification-cost.md` (LIN-2931), what lowers the cost of verifying and whether a verified artifact changes supervision. The next is commissioned: LIN-2930 (the frontier over four dates, projected, against each Harbour surface). Revise this document from those papers, never from the north star.

## Revision record

- v1, 2026-09-19: drafted from the planning conversation; not yet checked against a paper.
- v2, 2026-09-19: revised from the first paper. The gate narrowed to verification cost, budget kept as the operator's gate at rung 6; the big jump moved from 3-to-4 to handing over without watching; rung 3 marked unmeasured; the cohort reading and the social-adoption finding added; rung 5 re-sized; the rung-3 metric fixed to the tracker-movement form. Amendment proposed by the planning session and accepted by John Kershaw the same day.
- v3, 2026-09-19: revised from the second and third papers. The ceiling reading added (the tools set the ceiling in steps and the population fills the rungs below it); provenance named as the one handrail with a measured grip, and the claim that a verified artifact moves supervision withheld; the earlier-dates figures added; the permission-mode measurement corrected from "computable today" to "not recorded". Made by the planning session under the document's own rule and the authority John Kershaw gave it the same day to act on the papers' results.
