What is known about how developers adopt AI coding tools, and does it support the rung ladder?

Partly. The population has the shape a ladder predicts — as of April 2026 roughly 90% of developers use AI, 59% use an agent at work, about a fifth ever let one run unattended, and standing loops are not measured at all — and one vendor series shows the same individual climbing: auto-approve rises from about 20% of sessions for a new Claude Code user to over 40% for one with 750 sessions behind them. But three things in docs/ladder.md do not survive the evidence. The hardest step is not 3→4; agent use nearly doubled in eleven months and the stall is at handing over without watching. The reason is verification cost, not budget — budget barely appears in the population sources at all. And rung 3, the rung John's reading puts the majority on, is measured by nobody: no dated survey with a denominator asks whether developers save their prompts. The strongest disconfirming source, a study of tens of thousands of Microsoft engineers, finds the opposite of a climb: agent adoption spread socially, and prior experience with the lower rung made the higher one less likely to stick.

Findings

The ladder, as written. Six rungs (docs/ladder.md@0a60c5bf:15-20): 1, ask a question and copy the answer; 2, drive a session by hand, meandering; 3, bundle repeatable steps into saved prompts; 4, hand one bounded task to an agent; 5, ratify a passage of several tasks; 6, feed tasks into a standing loop. Each step is gated by "trust and budget". The doc's own load-bearing claims are that "the big jump is 3 to 4" (:32), that the majority of developers today sit at rung 3, and that "rungs 5 and 6 have one user today" (:38). It says plainly where its distribution comes from: "from John's own reading of the developers he works with" (:46). This paper is what that line asks for. Since it was written the same six steps have been accepted, unnumbered, into the normative layer — docs/north-star.md@0a60c5bf:17 now says Harbour "meets them at the saved prompt" and that "the measure is people per step and where they stop" — so the findings below bear on a rule, not only on a model.

Rung 1 is near-universal, and rung-1-only is a shrinking residual. Three independent surveys within five months agree: DORA, 90% of ~5,000 technology professionals using AI at work, fielded 13 June – 21 July 2025; JetBrains, 85% regularly using AI for coding, n=24,534, published October 2025; Stack Overflow, 84% using or planning to use and 51% of professionals using daily, n=49,009, fielded 29 May – 23 June 2025. None of them measures "chat and nothing else" directly, so the rung-1-only figure is a subtraction and should be read as one: JetBrains' January 2026 wave puts 74% on a specialised AI coding tool rather than a general chat interface against ~90% on some AI tool, implying ≈16% chat-only at January 2026 — a clean subtraction because both numbers come from the same wave, the same respondents, the same instrument. A second, weaker estimate for half a year earlier requires crossing instruments — Stack Overflow's 84% AI-use figure (mid-2025) against JetBrains' 62% coding-assistant figure (a different survey, a different question, a different population, published a month apart) — which would put the residual nearer 20–25%, but that comparison mixes two surveys' wordings and should not be read with the same confidence as the single-instrument figure. Estimate: rung 1 and above, 84–90%; rung-1-only, ≈16% at January 2026 (single-instrument), possibly 20–25% at mid-2025 (cross-instrument, weaker), and falling either way.

Rung 2 climbed steeply through 2026, but the three JetBrains figures are not one line. 62% relying on at least one AI coding assistant, agent or AI editor (published October 2025); 74% using a specialised coding tool (fielded January 2026); 90% using coding agents at work at least weekly and 68% daily, n>15,000 (fielded May–July 2026). The last of those uses "agent … in one form or another" as its category, which is broader than Stack Overflow's, and reading it as rung 4 would put agent use at 90% where the better-specified instrument says 59%. Mapped honestly it is a rung-2-and-above number. Estimate: rung 2 and above, 62% (late 2025) → 74% (January 2026) → ~90% (mid-2026).

Rung 3 is unmeasured. No source in this paper's population can say how many developers save their prompts. This is the rung the ladder puts the majority on, and it is the only rung for which searching produced no denominator in any phrasing — reusable prompt templates, prompt libraries, custom slash commands, CLAUDE.md/AGENTS.md adoption all return tutorials, vendor documentation and market-size projections, never a dated survey of developers. The one piece of real evidence that the behaviour exists is qualitative: an interview study of 17 experienced developers who use agents at least weekly (interviews July–August 2025, arXiv 2606.05391) names "a priori control" — configuring agent settings and custom instructions before delegating a task — as one of four forms of oversight work its participants routinely do. So rung 3 is real as a practice and absent as a population. John's reading is neither confirmed nor refuted, and the honest form of that finding is to leave it unbounded rather than to interpolate it between rungs 2 and 4.

Rung 4 roughly doubled in eleven months, on the same instrument. Stack Overflow 2025 (fielded May–June 2025): 31% use AI agents — 14.1% daily, 9% weekly, 7.8% monthly or less — with 17.4% planning to and 37.9% having no plans. Stack Overflow's April 2026 pulse (n=1,100, published 27 May 2026): 59% use agents at work at any frequency, stated explicitly as "compared to 31% in the 2025 Developer Survey". That is the strongest trend line available anywhere in this evidence, because it is one survey asking one question twice. The tool-side view agrees that rung 4 is where most agent use still sits: of 998,481 agentic API tool calls analysed by Anthropic (published 18 February 2026), 73% involve a human in the loop and 80% carry permission safeguards, with human involvement falling as tasks get harder (87% on simple tasks, 67% on complex). Estimate: rung 4 and above, ~31% (mid-2025) → ~59% (April 2026).

Rung 5 is where the population thins, and it is the one rung with published within-user evidence of a climb. Stack Overflow, April 2026: among agent users, 69% run a single agent, 17% multiple specialised agents and 16% multiple coordinated ones — 33% multi-agent — and 63% "rarely or never" let agents run entirely on autopilot, leaving ~37% who at least sometimes do, or about 20–22% of all developers. 60% block agents from making unapproved system changes and 68% say they prefer predictable single-agent setups. Against that cross-section, Anthropic's autonomy analysis (18 February 2026) measures the same users over time: auto-approve is used in about 20% of sessions by users with fewer than 50 sessions and over 40% by users with 750 or more, described as a gradual shift; experienced users also interrupt more often (9% of turns against 5%) because they stop approving each action up front and intervene when something goes wrong instead. Anthropic's June 2026 Economic Index adds a different kind of evidence, and it is not a second within-user series: Claude Code conversations score 0.37 points higher on a five-point autonomy scale than chat, a cross-surface comparison that holds task type constant rather than user identity — about two-thirds of the gap is "explained by the same tasks being executed with more delegation on Claude Code," the remaining third by a different mix of output types across the two surfaces, and the people on the two surfaces are not shown to be the same people, so part of the 0.37 may be exactly the surface-selection effect this paper's Limits warns about for vendor telemetry. The auto-approve-by-session-count curve above remains the paper's single published within-user evidence of a climb. Its separate January 2026 report tracks the population's share of directive conversations (minimal back-and-forth) over 2025, and that series is not a clean climb either: 27% in January 2025, rising to a peak of 39% in August 2025, then falling seven points to 32% by November 2025 — the report's own reading is that the August peak "overstated how quickly [automation] was materializing," though November still sat above the January baseline. So three different kinds of evidence sit in this paragraph, and they should not be read as agreeing on more than they do: one within-user series (auto-approve by session count) that climbs and holds; one population series (directive share) that rose and partly gave back within a year; and one cross-surface, task-controlled comparison (the 0.37-point gap) that is consistent with more delegation on Claude Code but cannot by itself say whether any individual climbed. Estimate: rung 5 and above, ~33–37% of agent users ≈ 20–22% of all developers (April 2026).

Rung 6 has no published population figure, and the ceilings around it are low. No survey in this paper's population has a category for feeding tasks into a standing loop. What exists bounds it from outside: 0.8% of agentic API tool calls are irreversible actions, and the 99.9th-percentile turn grew from under 25 minutes (October 2025) to over 45 minutes (January 2026) — a long tail measured in tens of minutes, not in standing operation. Anthropic's own 2026 trends report puts day-long autonomous runs in the future tense, offers a single seven-hour Rakuten run as the exemplar, and states that engineers use AI in roughly 60% of their work while being able to "fully delegate" only 0–20% of tasks. Estimate: no figure; the defensible reading is low single digits or less at April 2026. The ladder's "rungs 5 and 6 have one user today" (:38) is right about rung 6 and too pessimistic about rung 5.

The hardest transition is not 3 to 4. It is handing over without watching, and the reason named is the cost of verifying. docs/ladder.md@0a60c5bf:32 asserts "the big jump is 3 to 4 … at rung 4 something runs without them, and that is where trust breaks". The population does not stall there: agent use went 31% → 59% in eleven months, and the Microsoft rollout below shows the step happening across an organisation in a single quarter. Where the numbers fall off a cliff is between using an agent and not supervising it — 59% use one, ~22% ever let it run unattended, ~0 feed a standing loop. Every source that names a cause names verification rather than money: DORA's own language is that for an individual, friction "doesn't vanish so much as move: it shifts from manual grind to deciding and verifying," and that an unmanaged rise in change volume "could increase verification and coordination costs" — alongside its measured 30% reporting little or no trust in AI-generated code; Stack Overflow's trust falling while adoption rose, to 46% distrusting accuracy against 33% trusting and only 3.1% highly trusting in 2025, down from a combined 43.0% trusting (2.7% highly, 40.3% somewhat) among all respondents in 2024; 60% blocking unapproved changes and 68% preferring a single predictable agent; Anthropic's own ~60%-of-work against 0–20%-fully-delegated. Budget appears almost nowhere at the level of the individual developer, and organisational policy appears only as a configuration people choose for themselves. That asymmetry is a finding against the ladder's "trust and budget" framing (:15-20): on published evidence the gate is one thing, the price of checking the work, and it binds a rung later than the ladder expects.

The strongest disconfirming source

"Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI" (arXiv 2607.01418; adoption data 5 January – 29 April 2026 over tens of thousands of engineers, against a pre-period from 1 October 2025). It is the strongest because it is the only source with individual-level adoption and retention over a real population, rather than a cross-section of people describing themselves. Two findings cut against the ladder.

Adoption was social, not earned. Initial use was predicted overwhelmingly by exposure to other people: +216% higher odds of trying the agent where more than a quarter of skip-level peers had adopted, +82% where the manager used it. Tenure "barely mattered", and career stage produced only "a gentle gradient among individual contributors and nothing detectable among managers"; the authors' own summary is that "what an engineer does ... explains who adopts and retains far better than who an engineer is in the organization". If rung 4 were reached by accumulating trust through rungs 2 and 3, personal history should dominate. It did not; a colleague's adoption did.

Lower-rung experience predicted worse retention. Prior IDE Copilot use raised the odds of trying the CLI agent by 49–83% — and lowered retention by 12–15%, where retention means use on 5 of the 14 days from first use. The ladder predicts the opposite sign: the developer who has learned the sharp edges at rung 2 and saved prompts at rung 3 should be the one for whom rung 4 sticks. The sign is negative. The paper's own authors read this as a substitution effect: an engineer with a working IDE Copilot habit has "a familiar alternative to fall back on" when the CLI agent is harder to use, so a sustained new habit doesn't form — a mechanism compatible with a ladder whose rungs compete rather than build, not necessarily one that refutes the order outright. A second reading, equally consistent with their data, is the one this paper leans on: IDE-assistant fluency and CLI-agent delegation are different skills, so time on one doesn't transfer to retention on the other. Either reading leaves the sign negative, which a ladder where rungs accumulate does not predict.

A second disconfirming source fails the ladder in a different way. arXiv 2606.05391 (17 experienced developers using agents at least weekly, interviews July–August 2025) finds trust is contextual and task-dependent rather than accumulated: the same developer grants "unrestricted" autonomy on proof-of-concept work and reviews legacy-system work "more stringently than if it was a person", and confidence tracks the agent's demonstrated performance on that kind of problem rather than rising uniformly over time. On that reading a person does not occupy a rung at all; they pick one per task. Its participants also substitute plan-reading, test-passing and "eyeballing" for real review — "efficient, not perfect, oversight" — which is the verification tax being paid in a currency the ladder does not model.

A third source is not disconfirmation but a rival shape, and it is the most credible one available: DORA 2025 (fielded 13 June – 21 July 2025, ~5,000 respondents) deliberately retired its own decade-old elite-to-low performance ladder in favour of seven co-existing team archetypes, sized by its own cluster analysis: Harmonious high-achiever 20%, Pragmatic performers 20%, Constrained by process 17%, Stable and methodical 15%, The legacy bottleneck 11%, Foundational challenges 10%, High impact/low cadence 7% — and framed AI as an amplifier of existing capability rather than a thing climbed. A research programme with a long history of ladder-shaped models looked at AI adoption and chose not to use one. Likewise the grounded-theory study of 26 interviewees and 395 survey respondents (arXiv 2406.17325) reports no stage sequence at all, but a push-pull of motives and challenges — and finds that "creating a culture of sharing of AI best practices and tips" drives adoption, which is the same social mechanism the Microsoft study measured.

Method

Desk research on 19 September 2026, grounded at 0a60c5bf, in two passes: an original research pass (fourteen web searches, twelve direct fetches) and a revision pass after review, adding four further direct fetches — Anthropic's January 2026 Economic Index report, Stack Overflow's 2024 AI page, and the DORA 2025 report's full PDF (downloaded and extracted with pypdf, superseding the announcement pages the first pass had relied on for its figures) — made specifically to land two citations a review found short (see the header's model line for the dispatch items). Search terms covered, in order: Stack Overflow 2025 AI adoption and trust; DORA 2025; JetBrains agent adoption; Anthropic economic index and Claude Code automation; the METR trial; GitHub Octoverse 2025; 2026 agent adoption plateau, abandonment and churn; empirical and non-linear stage models of AI adoption; DORA's seven archetypes; depth-of-use data by organisation size; reusable prompt templates and prompt libraries; custom slash commands and CLAUDE.md/AGENTS.md usage; unattended and overnight agent operation; DORA fielding dates; and, in the revision pass, the exact wording of the Anthropic January 2026 delegation series and the DORA report's own language for verification cost. Sources were read at the primary URL wherever one existed; PDFs would not convert through the fetch tool and were extracted locally with pypdf.

The population is fifteen sources (listed in cites), classified before use into self-selected surveys (Stack Overflow ×3, JetBrains ×3, DORA), vendor telemetry (Anthropic ×3), vendor predictions (Anthropic's 2026 trends report), and independent studies (METR, arXiv 2607.01418, 2606.05391, 2406.17325). Each source's own categories were mapped onto the rungs one way only — from the source to the rung, never the reverse — and where a category did not map, as with JetBrains' broad mid-2026 "agent" wording, the paper says so rather than placing it.

Excluded, and why. SEO aggregator round-ups (digitalapplied.com, index.dev, masterofcode, prefactor.tech, uvik.net, sqmagazine, fungies.io, firstpagesage) restate the primary sources without adding data and frequently misattribute figures between them. Vendor maturity ladders (Salesforce's nine-stage "Agent Coding Maturity Curve", Augment Code's five-stage model, ELEKS) are frameworks with no population behind them; they are noted here as rival orderings and used for nothing else. Analyst projections such as "40% of agentic projects cancelled by 2027" are forecasts, not observations. Two sources were read and dropped for a different reason, on review: GitHub's Octoverse 2025 reports pull-request volume and repository counts, which describe platform activity rather than a developer's rung, and Anthropic's original April 2025 automation report (impact-software-development) turned out to supply no figure this paper's Findings does not already carry from a more specific Anthropic source. Neither is misquoted; neither earns a place in cites for a claim it does not carry.

Two searches returned nothing usable, and those empties are load-bearing. No survey data exists on saved or reusable prompts in any of three phrasings — that is what makes the rung-3 finding a gap rather than an oversight. And no population figure exists for unattended standing operation; the closest anything comes is Stack Overflow's 63% "rarely or never on autopilot", which bounds the complement rather than measuring the thing.

One deviation from the standard. docs/papers/standard.md@0a60c5bf:22-31 fixes a five-part order — Answer, Findings, Method, Limits, Next. This paper has six, with "The strongest disconfirming source" between Findings and Method, because LIN-2925 required that section and required it to be reached for honestly rather than folded into Findings where a supportive reading could absorb it. The deviation was put to John on the ticket and he chose the six-part form; the question of whether every paper should carry such a section is now a line in proposals.md. The paper also sits in harbour/ while being the archive's first paper about people outside Harbour, which stretches that directory's definition (:45-48); standard.md gains one line saying so rather than a third directory built for a single occupant.

Limits

Every survey here is self-selected, and they select for the AI-curious. Stack Overflow says so in its own methodology: respondents were recruited through Stack Overflow's own channels, and "highly-engaged users … were more likely to notice the prompts". JetBrains recruits from JetBrains' audience, DORA from Google Cloud's. A developer who uses no AI and reads no developer newsletter is the hardest person in this population to count, so every adoption figure in this paper is more likely too high than too low, and the rung-1-only residual is the estimate most exposed to that bias.

Three of the fifteen sources are vendor telemetry and one is a vendor prediction. The Anthropic autonomy and Economic Index figures — including the auto-approve curve that is this paper's best evidence for the ladder, and the directive-share series that partly reversed — measure Anthropic's own users, who are by construction further up it than the population; they are not neutral population data and must not be read as such. The 2026 Agentic Coding Trends Report is a predictions document; it is cited here only for the 60%-of-work against 0–20%-fully-delegated figure it attributes to internal research, and that figure is a self-report from Anthropic's own engineers with no published instrument.

Dates matter more than usual because the series moves fast. The oldest figures here are nine to fifteen months old at publication, and over an eleven-month window one of them doubled. Every number above carries its source's own date for that reason; a reader comparing two of them should check the dates before reading a difference as disagreement. The Stack Overflow 2026 pulse, which carries the rung-4, 5 and 6 estimates, has n=1,100 against the main survey's 49,009 — its confidence intervals are correspondingly wider, and the paper leans on its direction more than its precision.

The rung mapping is an interpretation and could be drawn differently. No source was built to answer this question, so every placement is this paper's judgement: "uses an agent" is not the same act as "hands over one bounded task", and a reader who thinks the ladder's rung 4 is narrower than Stack Overflow's "agents" should read the rung-4 estimate as an upper bound. The rung-3 gap is the sharpest consequence: because it is unmeasured, the paper cannot say whether the 90%-to-59% drop between rungs 2 and 4 happens at rung 3 or straight past it.

One source is read at one remove. The grounded-theory study (arXiv 2406.17325) was read at its abstract page after the PDF failed to convert through the fetch tool; its stage-sequence finding ("no sequential progression model... a push-pull of motives and challenges") is second-hand and a second edition should read the full text. The DORA figures no longer carry this caveat: the full 142-page report was downloaded directly (services.google.com/fh/files/misc/2025_state_of_ai_assisted_software_development.pdf) and read in this revision, and every DORA figure and quote in this paper — the 90% adoption, the 30% low-trust, the seven archetype percentages, the 13 June – 21 July 2025 fielding window, and the "shifts from manual grind to deciding and verifying" language — is quoted or paraphrased from that PDF, not from the shorter announcement pages this paper's first draft had relied on. Those announcement pages remain accurate for what they do state; they are simply not where the more granular figures live, which was the gap a review found.

Next

What Harbour could measure at each rung. The ladder proposes a funnel — people per rung, time on rung, and why they stop (docs/ladder.md@0a60c5bf:42) — and docs/north-star.md@0a60c5bf:17 now makes "people per step and where they stop" the measure. The first number it names, prompts copied before the first dispatch, is an activity count and can read green while the outcome is wrong. METR's trial is the evidence for exactly that failure: 16 experienced developers on 246 real tasks forecast a 24% speedup, reported a 20% speedup afterwards, and were measured 19% slower. Felt progress and real progress came apart by nearly 40 points in people with every reason to be accurate.

The fix is already half-written in the ladder: "the issue moved after the prompt was copied" (:30). That form tracks the outcome, because the tracker's state change is not something the user asserts. Stated explicitly: the tracker-movement form of the rung-3 metric is valid and should be built; prompts-copied alone is not. For rungs 4 to 6 Harbour already has externally witnessed signals and should prefer them to any self-reported rung — merged PR and discharged ledger at rung 4, rulings per landing report at rung 5, cost per verified task and forecast-against-actual at rung 6. And rung 5 has a published precedent for measuring a climb on behaviour rather than opinion: Anthropic's auto-approve share by session count. Harbour records permission mode on every run and can compute the same curve for its own users, which would make it the second organisation with within-user evidence on this question and the first outside a model vendor.

Two lines go to proposals.md: whether the standard should require a disconfirming section of every paper, and the second edition of this one.

The second edition should answer what this one could not. Read arXiv 2406.17325 in full for anything it says about saved prompts, since it was read here at its abstract only; and — the question this paper most wants answered — find or commission the first measurement of rung 3, since a rung that carries the product's positioning ("Harbour meets them at the saved prompt") currently rests on one person's reading of the developers he works with. A survey question Harbour could ask its own users would be a start, with the caveat that Harbour's users are selected for being at rung 3 already.

And a finding that wants a ruling rather than another paper. If the Microsoft study generalises — adoption is social, and lower-rung fluency slightly hurts higher-rung retention — then the rung-4 on-ramp is not primarily a trust-building exercise aimed at an individual who has climbed rungs 2 and 3. It is a distribution problem: the strongest predictor of a person trying an agent was that people around them already had. That is a different product shape from the one the ladder implies, and it is John's call, not this paper's.