---
title: Where does Harbour join the party — how fast is the frontier rung moving, and is each Harbour surface ahead of, level with, or behind it?
version: 1
date: 2026-09-19
authors: [Claude, John Kershaw]
model: claude-opus-5, Claude Code CLI, effort default (research item 5eea762e, plan); claude-sonnet-5, Claude Code CLI, effort default (implementation item 6e8b76ff-8552-443d-9974-a52bae5363b9, this session's primary-source re-verification and drafting)
grounded_at: 476604e8 (LinearViewer)
cites: [docs/ladder.md@476604e8:11, docs/ladder.md@476604e8:13, docs/ladder.md@476604e8:15, docs/ladder.md@476604e8:19-26, docs/ladder.md@476604e8:44, docs/ladder.md@476604e8:46, docs/ladder.md@476604e8:54, docs/north-star.md@476604e8:17, docs/papers/standard.md@476604e8:21, docs/papers/standard.md@476604e8:23-31, docs/papers/standard.md@476604e8:45-48, docs/papers/harbour/developer-adoption-ladder.md@476604e8, docs/papers/harbour/developer-adoption-ladder-earlier-dates.md@476604e8, docs/papers/harbour/what-lowers-the-verification-cost.md@476604e8, LIN-2925 (comment, 2026-09-19T10:11Z, ruling `lin2925-paper-order`), LIN-2930 (description, research comment of 2026-09-19T13:52Z, plan comment of 2026-09-19T13:59Z), code.claude.com/docs/en/agent-teams (read 2026-09-19), code.claude.com/docs/en/changelog (read 2026-09-19, v2.1.278), github.blog/changelog/2026-02-26-enterprise-ai-controls-agent-control-plane-now-generally-available/ (read 2026-09-19), claude.com/blog/cowork-for-enterprise (read 2026-09-19), claude.com/blog/claude-managed-agents (read 2026-09-19), services.google.com/fh/files/misc/dora-roi-of-ai-assisted-software-development-2026.pdf (v.2026.1, PDF creation metadata 2026-04-21T18:13:35+01:00, read 2026-09-19), dora.dev/ai/roi/report/ (states "Last updated: April 22, 2026", read 2026-09-19), infoq.com/news/2026/05/dora-roi-ai-assisted-dev-report/ (read 2026-09-19), metr.org/time-horizons/ (read 2026-09-19), survey.stackoverflow.co/2026/ (checked 2026-09-19, returns 404), stackoverflow.blog/2026/06/23/the-2026-developer-survey-is-now-open-for-human-developers-only/ (SO 2026 survey opening, read 2026-09-19), arxiv.org/abs/2607.01418 (read 2026-09-19 via LIN-2925), arxiv.org/abs/2606.05391 (read 2026-09-19 via LIN-2931), arxiv.org/abs/2406.17325 (read 2026-09-19 via LIN-2929), lib/brief.js@476604e8, lib/recap.js@476604e8, lib/render-observation.js@476604e8, lib/observation-sessions-store.js@476604e8, lib/prompt-templates.js@476604e8, lib/recommendation-facts.js@476604e8, routes/workspace-api-prompts.js@476604e8, routes/dispatch.js@476604e8, routes/passage-planner.js@476604e8, lib/render-passage-planner.js@476604e8, lib/scheduler.js@476604e8, lib/periodicals.js@476604e8, lib/pipeline-loops.js@476604e8]
---

# Where does Harbour join the party — how fast is the frontier rung moving, and is each Harbour surface ahead of, level with, or behind it?

"The frontier rung" names two different measurements, and keeping them apart is the answer's first move. The **ceiling frontier** — the highest rung any generally available tool offers — is a near-census, primary-verifiable series: it sat at rung 2 (on a strict general-availability reading) at September 2024, stepped to rung 4 on 25 September 2025 when GitHub's Copilot coding agent went GA, and had **not moved in the eleven months and 25 days since**, including at this paper's own cutoff of 19 September 2026, where Claude Code's agent teams remain disabled by default behind an environment variable. That series is what this paper projects: rung 4, bounds 4–5, at March 2027 (confidence medium-high), and rung 4–5 at September 2027 (confidence low, and no instrument supports rung 6 arriving by then). The **population frontier** — where developers actually sit — is a different, thinner series: its only same-instrument reading of rung 4 is Stack Overflow's agent-use question asked twice, 31% (mid-2025) rising to 59% (April 2026), and that is the *entire* evidential basis for every population-side claim below; it supports a bounded directional statement, not a number, and its next reading is already in the field and unpublished. Against those two series, one of Harbour's five surfaces — the bounded task with an evidence-verified merge, at rung 4 — sits level with the ceiling; passages (rung 5) and the standing loop (rung 6) sit ahead of it, by evidence that no rung-5 or rung-6 product is generally available anywhere; and the grounded next prompt (rung 3) sits behind it, because the ceiling has offered rung 4 to any new entrant since 25 September 2025 — a year and a half by the first projection date, two years by the second — while rung 3's own population has never been measured by any source at any of the four dates. Observation cannot be placed at all without a ruling: ladder v3 gives Harbour "Nothing yet" at rungs 1 and 2, and never names a rung for observation, which is not what the ticket that commissioned this paper assumed. This paper states the discrepancy, adopts a working reading, and hands the choice to John. It recommends no launch date.

## Findings

### 1. The frontier series and its projection

**One table, four dates, and the September 2024 cell keeps both of its readings — never averaged.**

| Date | Ceiling frontier rung | Gate named at that date | Same-instrument evidence |
|---|---|---|---|
| 30 Sep 2024 | **2** (strict GA) — *and, separately*, a rung-4 *analogue* reachable by a self-serve minority through `aider`, on no vendor's general terms | Trust, not budget (SSRN 4945566, 4,867 developers: less experienced adopters gained more) | `developer-adoption-ladder-earlier-dates.md`, §30 Sep 2024 |
| 30 Sep 2025 | **4**, GA five days before the cutoff — GitHub Copilot coding agent GA 25 Sep 2025 | DORA 2025: friction "shifts from manual grind to deciding and verifying" | `developer-adoption-ladder-earlier-dates.md`, §30 Sep 2025 |
| 31 Mar 2026 | **4** — unmoved; Claude Code Agent Teams shipped 5 Feb 2026 behind `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1` | Anthropic autonomy report: human-in-loop falls as task difficulty rises (87%→67%) | `developer-adoption-ladder-earlier-dates.md`, §31 Mar 2026 |
| **19 Sep 2026** | **4 — still unmoved** | Verification cost, unchanged since 31 Mar 2026 (`docs/ladder.md@476604e8:15`) | New to this paper: `code.claude.com/docs/en/agent-teams` (re-read this session) still opens with *"Agent teams are experimental and disabled by default. Enable them by setting `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`"*, and the changelog's latest entry is v2.1.278, dated 19 Sep 2026 — the documentation is current to the cutoff day itself |

**The ceiling has been rung 4 for 11 months and 25 days, and this session re-checked three candidate movers against it, adversarially, and none moves it.** Re-verified at primary this session, independently of the research comment that first found them:

- **GitHub's "Enterprise AI Controls & agent control plane," GA 26 Feb 2026** (`github.blog/changelog/2026-02-26-enterprise-ai-controls-agent-control-plane-now-generally-available/`): its own text is *"a suite of enterprise governance features designed to give GitHub Enterprise administrators deeper control and greater auditability around the use of AI controls and agents across their environments"* — administrator governance tooling, not a capability an ordinary developer uses to run a bounded multi-agent task.
- **Claude Cowork, GA 9 April 2026** (`claude.com/blog/cowork-for-enterprise`, primary, re-read this session): *"Claude Cowork is now generally available on all paid plans."* The same page states plainly it is not a developer surface: *"the vast majority of Claude Cowork usage comes from outside engineering teams,"* naming operations, marketing, finance and legal as the adopting functions. It does not offer a developer a ratified passage of coding tasks and is excluded from the ceiling series on that ground, now on a landed primary rather than the secondary press this paper's research phase had to rely on.
- **Anthropic Managed Agents, public beta 8 April 2026** (`claude.com/blog/claude-managed-agents`, primary, re-read this session): *"Managed Agents is available today in public beta on the Claude Platform,"* described as *"a suite of composable APIs for building and deploying cloud-hosted agents at scale."* Its own case studies (Notion, Rakuten, Asana) are companies building agent products on it, not developers handing it one bounded task with a click-to-merge. It is infrastructure for building the next rung's products, not a rung a developer occupies, and — like Cowork — is now landed on a primary rather than the secondary sourcing research left open.

**Confidence: high** that the ceiling has been rung 4 without interruption since 25 September 2025; **medium** for the September 2024 reading specifically, because the strict-GA and `aider`-analogue readings genuinely disagree and the paper carries both rather than picking one.

**Projection.** The ceiling series rests on a *class* of instrument — vendor release notes and changelogs, independently checked across three vendors (Anthropic, GitHub, OpenAI) at each date — which is its strength over anything below. **March 2027: rung 4, bounds 4 (floor — a ceiling does not fall) to 5 (upper, requiring agent teams or an equivalent multi-agent surface to leave preview). Confidence: medium-high.** The base rate cuts toward the upper bound, not against it: the feature has already sat in preview for 7½ months — already longer than the roughly four months the one comparable transition took (Copilot coding agent, public preview to GA) — so a GA release inside the six months to March 2027 would be overdue by that precedent, not premature; the medium-high confidence in the rung-4 floor rests on a single prior transition being a thin base rate to lean on, not on the wait arguing against GA. **September 2027: rung 4 to 5, genuinely open. Confidence: low.** No rung-6 candidate — a standing loop sold to an ordinary developer — is in preview anywhere in the evidence searched at this cutoff, so **rung 6 arriving by September 2027 is a defensible negative projection, not an absence of looking.**

**The population frontier cannot be given a number, and the honest form of that statement is this: all of the population-side projection rests on one instrument.** The only same-instrument rung-4 series across all four dates is Stack Overflow's agent-use question — 31% (n=49,009, fielded May–Jun 2025) to 59% (n=1,100, the April 2026 pulse, which states the comparison itself). Two points, one vendor, the second an order of magnitude smaller than the first. A naive linear extrapolation of that pair clears 100% before March 2027, which is self-evidently wrong and is shown here as a demonstration of why a point estimate is not available, not a projection to be believed. Bound every reading of it against `developer-adoption-ladder-earlier-dates.md`'s finding that the same "uses AI" question spans **42% to 97%** across instruments at one date — any movement smaller than that spread cannot be distinguished from instrument choice. The defensible form instead: **rung-4-and-above was about 59% in April 2026 on Stack Overflow's instrument, has not been re-measured on that instrument since, and is very unlikely to be lower at either projection date; it is bounded above by the 84–90% rung-1 figure. Confidence: medium for direction, low for any number.** **Rungs 3, 5 and 6 have no population figure at any of the four dates** — rung 3, the rung the product's own positioning names (`docs/north-star.md@476604e8:17`, "Harbour meets them at the saved prompt"), is unmeasured across three papers and roughly forty searches, and that absence is this paper's single most important qualification.

**The projection has a stated expiry, verified this session: `survey.stackoverflow.co/2026/` returns HTTP 404 at this cutoff.** The 2026 Developer Survey opened 23 June 2026; its 2025 predecessor was fielded in a comparable window and published 29 July 2025, roughly five weeks after fielding closed. The population-side numbers above are falsifiable by a single publication that may already be forthcoming, and this paper's population reading expires the day that survey's AI-agent question is published.

**METR's task-completion time horizons** (`metr.org/time-horizons/`, re-read this session) are cited here only as a leading indicator, with the caveat the primary page itself states, quoted directly: *"Measurements above 16 hrs are unreliable with our current task suite."* Per-model figures live only in an interactive graph, not in page text, so they are not quotable as a landed number here. This is a **capability** instrument — it bounds what could be built — not an adoption instrument, and it does not carry any projection in this paper.

### 2. What moved the frontier — and what the sources do not say moved it

**A tool shipping is the one cause with a clean date behind it.** The single ceiling step in the whole series is dated to Copilot coding agent's move from public preview (19 May 2025) to GA (25 Sep 2025); Claude Code's own GA (22 May 2025) is what first made the ladder's rung 2 available exactly as written. This is close to tautological — the ceiling is defined by availability — but it is the one causal claim with unambiguous dating.

**Verification cost is named by every population source that names a cause at all**, and DORA's second edition adds a mechanism neither prior paper had. `services.google.com/fh/files/misc/dora-roi-of-ai-assisted-software-development-2026.pdf` (v.2026.1) — downloaded and read in full this session; its PDF creation metadata is `2026-04-21T18:13:35+01:00` and the report's own page states "Last updated: April 22, 2026," so this paper resolves what the research phase had flagged as an unresolved date rather than leaving it open: **the report's own primary evidence puts its publication at 21–22 April 2026, a one-day gap plausibly explained by timezone or a publish-lag between file creation and page update, not a real disagreement. InfoQ's 11 May 2026 coverage is a press date, not a competing publication claim — that article states the report as "(2026.01)" and gives no release date of its own.** The report names a **J-Curve**, quoted directly: *"most organizations will encounter a J-Curve: a temporary productivity dip and period of instability associated with early adoption. Think of this dip as the tuition cost of transformation."* Its three named causes are *"the learning curve,"* *"the verification tax"* — quoted directly: *"Developers invest time reviewing generated code due to concerns about the trustworthiness of output"* — and *"pipeline adaptation."* This corrects a gap, not an error, in `what-lowers-the-verification-cost.md@476604e8`, which correctly found the phrase "verification tax" absent from DORA's **2025** report; it appears in DORA's **2026** report instead, which that paper does not cite. The fix belongs to that paper's own next edition, not to this one; this paper cites DORA 2026 for the J-Curve and moves on. Independently of DORA: DORA 2025's own language, *"friction doesn't vanish so much as move: it shifts from manual grind to deciding and verifying,"* Faros AI's telemetry across 22,000 developers (median time-in-review **+441.5%**), Sonar's finding that only 48% of respondents always check AI output before committing, and Harness's 81% reporting more time in review all name the same cost.

**Social spread is named by the strongest single-population source in the corpus.** arXiv 2607.01418 (tens of thousands of Microsoft engineers, adoption data January–April 2026) found trying the agent predicted by peer exposure — **+216%** odds where more than a quarter of skip-level peers had adopted, **+82%** where the manager had — with tenure "barely" mattering.

**Two causal claims the sources do not support, stated as absences rather than softened:** **no source shows a price change moving anyone's rung** — budget is named as an individual-level gate by no source at any of the four dates (`docs/ladder.md@476604e8:15`); the one pricing datum found (a coding-agent subscription price rise, secondary-sourced) carries no adoption measurement. **No source shows verification getting cheaper and a rung opening as a result** — `what-lowers-the-verification-cost.md@476604e8` found the only candidate with a controlled positive result is provenance, not tests or CI, on a web-agent study of fourteen participants, and that no published study shows a verified artifact changes how much a developer supervises; two controlled studies that manipulated the signal directly found behaviour barely moved. This paper does not assert a causal chain the source material does not support.

### 3. Harbour's surfaces against the frontier

All five surfaces named below exist in the codebase at HEAD (verified this session): `lib/brief.js`, `lib/recap.js`, `lib/render-observation.js`, `lib/observation-sessions-store.js` (observation); `lib/prompt-templates.js`, `lib/recommendation-facts.js`, `routes/workspace-api-prompts.js` (the grounded next prompt); `routes/dispatch.js` and the review ledger (one bounded task); `routes/passage-planner.js`, `lib/render-passage-planner.js` (passages); `lib/scheduler.js`, `lib/periodicals.js`, `lib/pipeline-loops.js` (the standing loop). Ladder v3's "Harbour gives" column (`docs/ladder.md@476604e8:19-26`) is the authoritative rung mapping; it is reused here, not re-derived.

| Harbour surface | Rung (ladder v3) | vs. ceiling at Mar 2027 (rung 4, bound 4–5) | vs. ceiling at Sep 2027 (rung 4–5, low confidence) |
|---|---|---|---|
| Observation (briefs, recaps, Observation view) | **Not on any rung** — ladder v3 says "Nothing yet" at rungs 1–2 | see the discrepancy below — cannot be placed without a ruling | same |
| The grounded next prompt | 3 | **Behind by one rung** — evidence: the ceiling has stood at rung 4 since Sep 2025, so a developer entering after that date never sees a world where rung 3 is the top of what tools offer; and rung 3's population is unmeasured at all four dates, so no population figure can offset that reading | Behind by one to two |
| One bounded task, evidence-verified merge (dispatch, PR, CI, ledger, one approval click) | 4 | **Level with the ceiling** — evidence: this is the only surface whose rung equals the highest rung any generally available tool offers under the base case | Level, or behind by one if the ceiling steps to 5 |
| Passages (legs, task budgets, landing report, rulings) | 5 | **Ahead of the ceiling by one rung** — evidence: no rung-5-shaped product (agent teams, Agent HQ) is generally available; each sits in preview or behind a flag | Ahead by one, or level if a multi-agent surface reaches GA |
| The standing loop (always-on, forecast, KPIs) | 6 | **Ahead by two rungs** — evidence: no rung-6 candidate is even in preview anywhere in the evidence searched | Ahead by one to two; no instrument supports rung 6 arriving by this date |

**"Ahead" and "behind" are stated above with the evidence that supports them, not the bare word, per the paper's own requirement.** The sharpest of the five readings is the grounded next prompt's: it is behind not because a population figure says so — there isn't one for rung 3 at any date — but because the ceiling series itself, which is valid and near-census, has stood at rung 4 since 25 September 2025 — a year and a half by the first projection date, two years by the second. `docs/ladder.md@476604e8:13`: *"a cohort's entry rung is the ceiling on its start date."* A developer who starts in 2027 starts at the ceiling that already exists, not at the rung the product is positioned to meet them.

**The discrepancy this paper poses rather than resolves.** The ticket that commissioned this paper carries John's own framing: *"it first steps in as observation — watching tasks, getting analysis."* Ladder v3 does not say this. It gives Harbour "Nothing yet" at rungs 1 and 2, and names no rung for observation anywhere in its rung table. Three readings are defensible, and they give different answers to whether observation arrives too early:

1. **Observation is rung-3 evidence infrastructure**, not a rung-1/2 offer. `docs/ladder.md@476604e8:44`: *"Rung-3 evidence is free. Harbour reads the tracker, so it can see whether an issue moved after a prompt was copied without asking the user anything."* Observation is the mechanism that reading names. On this view observation sits level with the grounded next prompt, at rung 3, and behind the ceiling by one.
2. **Observation is an operator surface**, read at rungs 5–6 to ratify passages and build Harbour itself (`docs/ladder.md@476604e8:54`). On this view observation is ahead of the ceiling.
3. **Observation is rung-agnostic** — it watches whichever rung a person is on, which is arguably what makes it offerable to anyone regardless of where the frontier sits.

**This paper adopts reading (1) as its working mapping for the table above, because it is the only one the ladder's own text supports directly, states the other two, and hands the choice to John rather than deciding it or editing the ladder.** The ladder is revised from papers; this paper is an input to that revision, not the mechanism of it, and no edit to `docs/ladder.md` is made in this PR.

### 4. The cohort question

**Ladder v3 already states the mechanism; this paper does not re-derive it.** `docs/ladder.md@476604e8:11`, `:13`: each cohort's entry rung is the ceiling on its own start date, and most of the movement between ceilings is people filling rungs that already exist rather than new rungs appearing. The evidence for cohort-skipping specifically is indirect and negative-signed: arXiv 2607.01418 found prior IDE-assistant use **lowered** CLI-agent retention by 12–15%, against a ladder that predicts the opposite sign for someone who has climbed the lower rungs already; SSRN 4945566 (4,867 developers, three firm-run field experiments, 2024) found less-experienced developers adopted more and gained more, the same sign two years earlier. **No source at any of the four dates tracks the same individuals across rungs over time.** The nearest available evidence, Anthropic's auto-approve-by-session-count curve, is within-user but cannot distinguish a climb from a skip. **Confidence: medium for the mechanism, low for any magnitude.**

**What this does to a product that meets people at rung 3 is a question this paper poses and does not answer.** If the ceiling has offered rung 4 generally since 25 September 2025 — a year and a half by the first projection date, two years by the second — a developer entering in 2027 has never seen a world where rung 3 was the top of what tools offered — the saved prompt may then be a technique some entrants skip rather than a rung they pass through. Against that: rung 3 has never been measured, at any of the four dates, by any source this paper or its two predecessors found, so "most developers are at rung 3" cannot be used either to confirm or to reassure. Both facts point the same way, and neither is a measurement.

## The disconfirming case

This paper's guard is the case that timing does not matter — that developers adopt whatever is in front of them regardless of rung, or that the frontier is not one line but several populations moving separately.

**"Several populations, not one line": well supported, by four real sources.** **DORA's 2025 report** (n≈5,000, fielded 13 Jun–21 Jul 2025) retired its own decade-old elite-to-low performance ladder for **seven co-existing team archetypes** sized by cluster analysis (Harmonious high-achiever 20%, Pragmatic performers 20%, Constrained by process 17%, Stable and methodical 15%, The legacy bottleneck 11%, Foundational challenges 10%, High impact/low cadence 7%), framing AI as an amplifier of existing capability rather than a rung climbed — this is the shape the ticket named for testing, tested here directly rather than mentioned in passing, and it is still the frame DORA's own 2026 ROI report carries forward. **arXiv 2406.17325** (26 interviewees, 395 survey respondents) explicitly tested and rejected a staged-adoption model: a prior preliminary stage model (Russo 2023) had, in its words, "3 of 7 hypotheses ... not supported upon further analysis," and it develops instead a non-sequential push-pull framework of individual and organisational motives and challenges; it never mentions saved or reusable prompts as a behaviour. **arXiv 2606.05391** (17 experienced developers) finds trust **contextual and task-dependent, not accumulated** — the same developer grants "unrestricted" autonomy on proof-of-concept work while reviewing legacy-system work "more stringently than if it was a person." On that reading a person does not occupy a rung; they pick one per task, which this paper treats as the strongest of the four because it cuts at the ladder's central assumption directly. **arXiv 2607.01418** is, independently of its cohort-sign finding above, itself evidence that adoption is a distribution phenomenon rather than an earned climb.

**"Timing does not matter — developers adopt whatever is in front of them": no source was found that states this directly, and this paper says so plainly rather than manufacturing support.** The searches run (by the research phase this paper reuses, and re-checked for anything published since) covered general-availability timing, adoption-timing framing, and abandonment/plateau framing, in addition to the disconfirming-interpretation searches listed in Method; none returned a source arguing timing is irrelevant. The closest available support is indirect: the Microsoft rollout's +216% peer-exposure effect means what was *socially* in front of a person dominated who adopted, which is a claim about proximity, not about readiness — adjacent to, but short of, "timing does not matter."

**A fifth source, found by this paper's own research phase and cited by neither prior paper, reframes the question rather than answering either side of the guard.** DORA 2026's J-Curve (§2 above) states that the cost of adoption is a temporary dip paid by every adopter — the learning curve, the verification tax, pipeline adaptation — rather than a cost that falls only on whoever adopts at the wrong rung or the wrong time. On that reading, "where does Harbour join the party" is not fully answered by a rung comparison at all: the sharper question DORA's evidence raises is not *when* Harbour joins, but *whether Harbour shortens the dip* for whoever it meets. This paper states that reframing and leaves it as a reframing, not an answer; it is not this paper's to resolve.

**No source in any of the three papers splits the population by rung and cohort simultaneously — which is exactly what reading "the frontier" as a single line over four dates assumes exists.** The DORA archetypes split teams seven ways; the Microsoft study splits by peer exposure; arXiv 2606.05391 splits by task, within one person. None of the three splits by rung and entry-cohort at once. Stated as a limit on the whole exercise, not resolved here.

## Method

**Reused, not re-derived.** This paper reuses `developer-adoption-ladder.md@476604e8`'s instrument-to-rung mappings for recurring instruments (Stack Overflow's agent-use question, Anthropic's auto-approve series) and its Microsoft-rollout finding; `developer-adoption-ladder-earlier-dates.md@476604e8`'s three earlier frontier cells verbatim, including the two-reading September 2024 cell, its same-instrument/cross-instrument separation as a structural rule, and its 42–97% cross-instrument bound; and `what-lowers-the-verification-cost.md@476604e8`'s causal findings on verification cost. It does not duplicate the first paper's fifteen-source population inventory or re-run the second paper's per-date search.

**The fourth date is this paper's own new fact.** Neither prior paper states a September 2026 frontier rung — `developer-adoption-ladder.md` was never asked for one. It was established from the primary sources named in §1 above, all re-verified directly by this implementation session (not merely inherited from the research phase's earlier reads): `code.claude.com/docs/en/agent-teams`, `code.claude.com/docs/en/changelog`, and the three candidate-mover checks (GitHub's agent control plane, Cowork, Managed Agents), each read at a landable primary URL.

**Three primary-source checks the research phase left open were closed in this implementation session, not deferred:**

1. **DORA 2026 ROI report's publication date.** Resolved by downloading the primary PDF directly (`services.google.com/fh/files/misc/dora-roi-of-ai-assisted-software-development-2026.pdf`) and reading its embedded creation metadata (`2026-04-21T18:13:35+01:00`) against the page's own "Last updated: April 22, 2026" — a one-day gap, not the unresolved multi-week disagreement the research phase's secondary sourcing had suggested; InfoQ's 11 May 2026 date is a coverage date, stated by InfoQ itself without a competing publication claim.
2. **Primaries for Cowork GA and Managed Agents.** Both were secondary-sourced only when research left off. Both are now landed at primary vendor URLs (`claude.com/blog/cowork-for-enterprise`, `claude.com/blog/claude-managed-agents`), so both remain in the rejected-ceiling-movers discussion in §2 — the standard's citation-landability rule (`docs/papers/standard.md@476604e8:21`) is satisfied rather than worked around.
3. **METR's time-horizon figures.** Confirmed directly against the primary page: per-model figures remain in an interactive graph, not page text, and the 16-hour reliability caveat was re-read and is quoted verbatim in §1. METR is cited only as a leading capability indicator and carries no projection.

**Ladder citation discipline.** Every `docs/ladder.md` citation in this paper is at `476604e8` — the HEAD commit at the time research, planning and writing were each grounded, re-checked immediately before this paper was drafted (`git log --since="2026-09-19T10:32:37Z" -- docs/ladder.md docs/north-star.md docs/papers/`, unchanged) — never at the prior papers' own citation SHAs (`0a60c5bf` for the first paper, `b5a4f77f` for the second and third).

**Six-part order, on a standing ruling, not re-raised.** `docs/papers/standard.md@476604e8:23-31` fixes five parts (Answer, Findings, Method, Limits, Next). This paper carries six, with the disconfirming case between Findings and Method, under ruling `lin2925-paper-order` (LIN-2925, comment of 10:11 UTC 19 September 2026), which both prerequisite papers already carry forward. `docs/papers/standard.md@476604e8:45-48` already widens `harbour/` to cover papers about the people Harbour serves, so no location deviation is needed either.

**Search terms for the two negative results this paper states rather than papers over.** For "timing does not matter": adoption-timing irrelevance, rung-agnostic adoption, AI tool adoption regardless of maturity stage, developer adoption independent of tool readiness. For post-cutoff evidence: developer AI agent adoption survey published October 2026; AI coding agent news published "September 2026" developer adoption report; DORA State of AI-assisted Software Development 2026 second edition release; Stack Overflow developer survey 2026 results published; coding agent unattended autonomous "generally available" June–August 2026; "agent teams" Claude Code exits research preview. None returned a source dated on or after 19 September 2026 — the cutoff is today, so that set is empty and could not have been otherwise, and this paper states that rather than dressing it up. One aggregator claim, "71% of professional developers report using an AI coding agent at least daily according to the Stack Overflow Developer Survey 2026," was checked and excluded: no such survey exists (`survey.stackoverflow.co/2026/` returns 404), the same failure mode the earlier-dates paper documented for a re-badged "KPMG Q1 2026" figure.

## Limits

**The published-evidence layer cannot be closed.** Every class of source in this paper — frontier-gate causes, projection instruments, disconfirming interpretations — was bounded by search, not by enumeration; a negative literature result (no fifth instrument, no direct "timing doesn't matter" source) is a bound on what was found, not a proof that nothing exists. The in-repository layers (the ladder's versions, the two prerequisite papers, the codebase's surface files) were bounded exhaustively by `git log` and file existence checks and are exhaustive.

**Rung 3 is unmeasured at all four dates, by every source three papers and roughly forty searches have found.** This is the rung the product is positioned at (`docs/north-star.md@476604e8:17`) and the rung this paper's surfaces table can least defend with a number.

**The entire population-side projection rests on one instrument, whose next reading is already in the field.** Stack Overflow's 31%→59% rung-4 series is two points from one vendor, the second an order of magnitude smaller in sample than the first; no other same-instrument rung-4 series exists anywhere in the four-date span. The 2026 Developer Survey opened 23 June 2026 and its results are unpublished at this cutoff — this paper's population-side reading expires the day that survey's agent-use question is published, and could be strengthened or overturned by it.

**The September 2024 ceiling boundary is a judgement call, shown rather than resolved.** "Generally available" is the rule applied identically at all four dates, but it produces two different answers at the earliest one — rung 2 on a strict reading, a rung-4 analogue via a self-serve tool on a looser one — and this paper, like its predecessor, carries both rather than picking one.

**No source splits the population by rung and cohort at once**, which is what treating "the frontier" as a single line across four dates implicitly assumes is measurable. DORA's archetypes, the Microsoft peer-exposure study, and the per-task-trust study each split the population a different way; none does both at once, and this paper cannot manufacture that split from what exists.

**The ceiling-to-population causal link is correlational, not demonstrated.** Nothing in the evidence dates an individual developer's climb to a specific product's GA date; the one within-user series available, Anthropic's auto-approve curve, is confounded by capability gains over the same window and cannot separate "trusted the artifact more" from "the agent got better."

## Next

**One line is required by the standard and by this ticket's own acceptance criteria; it goes here and to `docs/papers/proposals.md` in this PR:** build the entry-rung measurement Harbour can run on its own users — which rung a new user's first successful action lands on, read from tracker movement and the date of their first evidence-verified merge, against `docs/ladder.md@476604e8:19-26`'s rung table. This tests the cohort claim in §4 on evidence Harbour already holds, rather than waiting on another population survey Harbour does not control.

**A second candidate, not added to `proposals.md` in this PR because it belongs to another paper's line of work:** `what-lowers-the-verification-cost.md`'s next edition should absorb the DORA 2026 "verification tax" finding this paper surfaced in §2, since that paper's own search for the phrase, correctly, came back empty against the 2025 report only.

**And a question that is John's, not this paper's, restated once more so it is not lost in the middle of Findings:** which of the three readings in §3 is right for observation — rung-3 evidence infrastructure, an operator surface at rungs 5–6, or rung-agnostic — changes whether observation is ahead of, behind, or unplaceable against the frontier, and this paper cannot resolve it from `docs/ladder.md`'s current text alone.
