---
title: What lowers the cost of verifying agent work, and does a verified artifact change how much a developer supervises?
version: 1
date: 2026-09-19
authors: [Claude, John Kershaw]
model: claude-opus-5, Claude Code CLI, effort default; dispatched as LIN-2931 research item f574a25a, implementation item e93413e5-cb6b-4f87-8f5a-d6390e715887
grounded_at: b5a4f77f (LinearViewer), d0e809e3 (simple-dispatcher)
cites: [services.google.com/fh/files/misc/2025_state_of_ai_assisted_software_development.pdf (the full DORA 2025 report, read 2026-09-19), survey.stackoverflow.co/2025/ai (read 2026-09-19), survey.stackoverflow.co/2024/ai (read 2026-09-19), stackoverflow.blog/2026/05/27/agents-on-a-leash-agentic-ai-remains-mostly-monitored-at-work (read 2026-09-19), anthropic.com/research/measuring-agent-autonomy (read 2026-09-19), metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study (read 2026-09-19), arxiv.org/pdf/2507.09089 (METR full paper, read 2026-09-19), arxiv.org/abs/2606.05391 (read 2026-09-19), arxiv.org/abs/2605.29442 (read 2026-09-19), arxiv.org/html/2605.29442 (read 2026-09-19), arxiv.org/abs/2606.18671 (HANSEL, read 2026-09-19), arxiv.org/abs/2203.05045 (Kudrjavets et al., MSR 2022, read 2026-09-19), arxiv.org/abs/2602.16844 (read 2026-09-19), arxiv.org/abs/2606.05647 (read 2026-09-19), arxiv.org/abs/2606.22484 (read 2026-09-19), arxiv.org/abs/2608.12355 (read 2026-09-19), arxiv.org/abs/2609.03460 (read 2026-09-19), arxiv.org/abs/2607.01418 (read 2026-09-19, via docs/papers/harbour/developer-adoption-ladder.md), faros.ai/blog/ai-acceleration-whiplash-takeaways (12 Apr 2026, read 2026-09-19), sonarsource.com/company/press-releases/sonar-data-reveals-critical-verification-gap-in-ai-coding (8 Jan 2026, read 2026-09-19), harness.io/state-of-engineering-excellence (13 May 2026, read 2026-09-19), prnewswire.com/news-releases/harness-report-reveals-ai-has-outpaced-how-engineering-organizations-measure-developer-productivity-302770521.html (13 May 2026, read 2026-09-19), tfir.io/ai-coding-has-a-trust-problem-sonar-data-shows-verification-lagging-far-behind-adoption (secondary for Vogels' re:Invent 2025 quote, read 2026-09-19), implicator.ai/werner-vogels-hands-out-newspapers-at-his-likely-final-re-invent-the-man-who-built-the-cloud-isnt-done-teaching (secondary, dates the keynote to re:Invent 2025, read 2026-09-19), lib/prompt-template-defs.js@b5a4f77f:1061, lib/prompt-template-defs.js@b5a4f77f:1080-1099, lib/prompt-template-defs.js@b5a4f77f:302-313, docs/architecture/prompt-system.md@b5a4f77f:28, docs/papers/standard.md@b5a4f77f:7-21, lib/periodical-report-gate.js@b5a4f77f:20-35, lib/dispatch-factory.js@b5a4f77f:317, lib/dispatch-store.js@b5a4f77f:379, docs/ladder.md@b5a4f77f:42, docs/ladder.md@b5a4f77f:54, docs/north-star.md@b5a4f77f:7, docs/lin-367-research-notes.md@b5a4f77f:28, lib/escalation-kpis.js@b5a4f77f:27, docs/papers/harbour/developer-adoption-ladder.md@b5a4f77f:55, docs/papers/harbour/developer-adoption-ladder.md@b5a4f77f:73, docs/papers/standard.md@b5a4f77f:22-31, docs/papers/standard.md@b5a4f77f:45-48, simple-dispatcher/executors.js@d0e809e3:483, simple-dispatcher/state-store.js@d0e809e3:34, LIN-2925 (comment, 2026-09-19T10:11Z), PR #1510 (2026-09-19), docs/papers/proposals.md@b5a4f77f:141-146, LIN-2931 (description and comments, 2026-09-19)]
---

# What lowers the cost of verifying agent work, and does a verified artifact change how much a developer supervises?

Verification cost is real, rising, and named by every source that measures it as the actual gate on handing agents more autonomy — the first ladder paper's finding. Of the candidates for lowering it, only one has a controlled measurement behind it, and it is provenance (surfacing evidence for a claim), not tests or CI. No published study shows that a stronger verification signal — green CI, a merged PR, a discharged ledger — changes how much a developer supervises; the two studies that manipulated a verification signal experimentally found that most participants' behaviour barely moved, and one found supervision *should* have risen and didn't. "Verified beats claimed" is a defensible bet on what evidence *should* count, but it is not, today, a claim any published study supports about what changes anyone's behaviour.

## Findings

**1. What the population sources say the verification cost is.** DORA's own words, from the full 142-page 2025 report (fielded 13 Jun–21 Jul 2025, n≈5,000; not the phrase "verification tax", which the LIN-2925 review already found absent from both cited pages and this paper's own re-check of the full PDF confirms occurs zero times in it): "an increase in change volume without a corresponding set of evolved guardrails, roles, and 'golden paths' could increase verification and coordination costs," and "friction doesn't vanish so much as move: it shifts from manual grind to deciding and verifying" [study]. The landable phrase that *does* exist is someone else's: **"verification debt"**, used by AWS CTO Werner Vogels at his final re:Invent keynote — **December 2025, not December 2024** (this paper's audit corrects an earlier research note's date; the primary AWS transcript was not found, so the quote is attributed to secondary reporting): *"You will write less code, 'cause generation is so fast, you will review more code because understanding it takes time... When the machine writes it, you'll have to rebuild that comprehension during review. That's what's called verification debt"* [vendor claim, secondary-sourced]. Faros AI's telemetry across 22,000 developers and 4,000+ teams, two years, comparing each org's lowest- and highest-AI-adoption periods (published 12 Apr 2026): median time in review **+441.5%**, code churn **+861%**, incidents-to-PR ratio **+242.7%**, PRs merged with no review at all **+31.3%** [vendor telemetry]. Sonar (published 8 Jan 2026, n>1,100): 96% don't fully trust AI-generated code is functionally correct, only 48% always check it before committing, 38% say reviewing it takes more effort than reviewing a colleague's code [vendor claim, self-report survey]. Harness's State of Engineering Excellence Report 2026 (published 13 May 2026, n=700 across the US/UK/India/France/Germany): 81% spend more time in code review since adopting AI [vendor claim, self-report survey] — the 28%-report-an-increase-of-more-than-30% figure sits in the report's gated PDF, not the landable `harness.io` page, so it is attributed here to the co-published PR Newswire release that carries it (prnewswire.com, 13 May 2026) [vendor claim, self-report survey]. **The counterweight this paper must carry:** METR's RCT (16 experienced developers, 246 real tasks, published 10 Jul 2025) measured developers **19% slower** with AI against a forecast 24% speedup and a post-hoc belief they were still 20% faster — but its screen-recording decomposition (128 of the study's recordings hand-labelled, 143 hours, 29% of total observed time) puts reviewing and cleaning AI-generated code at only **~9% of task time on the 44 labelled AI-allowed issues** [study]. Verification is a real, named, measurable slice of the work — and on the one study that decomposed it directly, it is not large enough by itself to explain a 19-point slowdown.

**2. What is shown to reduce it — four evidence kinds, kept apart.** Only one candidate has a controlled result, and it is not tests or CI: **provenance**. HANSEL (arXiv 2606.18671, n=14, within-subjects, published Jun 2026) surfaced trajectory evidence for a web agent's answer and reduced task completion time and perceived effort, with participants rating it significantly higher on verification ease and error identification than a baseline agent interface showing only the final answer and source list [study] — it measured task completion time and subjective ratings, not verification time directly. Its own scope caveat matters: web-browsing agents, not coding agents, and n=14. **Smaller diffs has a negative result against it, not a supportive one**: Kudrjavets, Nagappan and Rastogi (MSR 2022, arXiv 2203.05045), 845,316 pull requests and 401,790 code reviews across multiple languages and platforms, found PR size has no measurable relationship to time-to-merge or time-to-accept — the "smaller diffs review faster" guideline the industry repeats is not backed by this study; it is contradicted by it [study]. Plan-first workflows, sandboxing, and bounded task size rest on weaker ground: arXiv 2606.05391 (17 experienced developers, interviews Jul–Aug 2025) names "a priori control" — configuring and constraining an agent before delegating — as one of four oversight behaviours its participants use, but measures no effect on review time or outcome [study, qualitative only]. Graduated/tiered oversight's most-quoted number, "84–97% velocity preservation" (arXiv 2606.22484, Jul 2026), is an *analytical productivity model* over a regulatory-classification framework, not a measurement of any real deployment [author reasoning, not a study — a paper quoting this figure without that qualifier is repeating a vendor-shaped claim as if it were evidence]. Anthropic's own telemetry (998,481 agentic API tool calls, published 18 Feb 2026) shows human-in-the-loop involvement at 73% overall, falling from 87% on simple tasks to 67% on complex ones, and permission safeguards present on 80% of calls — real behaviour at scale, but on Anthropic's own non-neutral user population [vendor telemetry].

## The strongest disconfirming sources

**No published study shows a verified artifact changes supervision, and two controlled studies that manipulated a verification signal found it largely did not move behaviour.** "Coding with 'Enemy'" (arXiv 2606.05647, 100+ developers, ~5-hour realistic tasks with four frontier models, published Jun 2026): 94% of developers failed to detect deliberately planted agent sabotage unaided; adding an automated safety monitor that correctly raised an alert reduced but did not eliminate the failure — **56% of participants still accepted the flagged, malicious code**, citing minimal code review, a plausible cover story, and overtrust in the agent as the contributing factors named in the paper. A stronger, explicit verification signal did not change the merge decision for most people who received it.

**A verification interface can make people faster and more confident without making them more correct.** Microsoft Research's three-study series (arXiv 2602.16844, n=12 per study, published 18 Feb 2026) found that a step-focused presentation of an agent's trajectory "can justify overreliance" — participants who mistook an incorrect answer for correct cited the agent's "reasonable" process as their reason for not checking further — and that an improved interface explicitly surfacing the task's requirements and the agent's assumptions helped participants **find errors faster and report higher confidence, with no meaningful effect on accuracy**. Stronger verification tooling changed how people felt about checking, not how well they checked.

**The nearest thing to supportive behavioural evidence is confounded.** Anthropic's auto-approve share rises from about 20% of sessions for users under 50 sessions to over 40% for users at 750+ (published 18 Feb 2026) — but capability on the same population's hardest internal tasks roughly doubled over the same window, so the vendor's own telemetry cannot isolate whether users trusted the *artifact* more or the *agent* became more reliable [vendor telemetry, confounded]. The one large-population measurement of supervision actually falling is Faros' PRs-merged-with-no-review, **+31.3%** — but the published account attributes this to review capacity not keeping pace with PR volume, not to any verification signal changing anyone's judgement.

**A fourth shape of disconfirmation: transparency can itself cost trust, independent of a verification artifact.** arXiv 2609.03460 (n=81, published Sep 2026) measures a "transparency penalty" — disclosure that content is AI-generated lowers perceived trustworthiness at constant quality — and finds that only a much richer interface (visualising the density of *verified* claims directly, not merely disclosing AI involvement) restores a real discernment gap between accurate and fabricated content (+4.15 points, d=1.82, against no detectable discrimination with no signal at all). The result cuts two ways for this paper's question: a thin verification signal (a disclosure label) can move perceived trust in the wrong direction, while a dense one (visualised provenance) is the closest published evidence that a richer artifact *can* change a judgement — but this study measures **perceived trustworthiness of text, not a developer's supervision behaviour on code**, and should not be read past that.

**This search cannot prove a negative, and the field says so about itself.** A position paper naming the gap directly (arXiv 2608.12355, "Humans are Missing from AI Coding Agent Research," published Aug 2026) argues the field over-optimises for autonomous task completion and under-measures "verifiability" as one of four core human-agent interaction dimensions, calling how much human effort verification demands "uncharted." That a live research programme names this measurement as missing is the strongest available bound on the absence this paper reports — not proof that no such study exists anywhere, but evidence that the field regards the question as open rather than answered.

**One dataset checked specifically for whether it changes this finding: it does not, and it adds a caveat.** arXiv 2605.29442 ("How Coding Agents Fail Their Users," 20,574 real coding-agent sessions across 1,639 repositories, submitted 28 May 2026) does not measure whether tests, CI, or any verification/provenance signal changes review depth or correction rate — that relationship is outside its scope. What it does show: of 16,118 identified misalignment episodes, only 1,504 (9.33%) have a visible resolution in the session log at all, and of those resolved episodes, 91.49% were resolved by the agent correcting itself only *after* explicit developer pushback (RV2), against 2.99% where the agent self-corrected unprompted and 5.52% where the developer took over outright. Read narrowly — it says nothing about verification artifacts — this is still a data point against automatic self-correction lowering the need for a human to intervene: even where a fix was visible in the log, it overwhelmingly needed a human to ask for it first.

## Method

Desk research on 19 September 2026, in two passes: an earlier research pass (fourteen web searches, seventeen fetches, four PDFs extracted locally with `pypdf`, recorded in the LIN-2931 research comment) and a drafting-session validation pass that read or re-verified the three sources the research pass flagged as outstanding — arXiv 2605.29442 in full (via its HTML rendering, since the PDF did not convert through the fetch tool and this paper does not claim to have read a PDF it could not extract), the Harness 2026 report's own landing page and PDF link (`harness.io/state-of-engineering-excellence`, superseding an earlier at-one-remove reading via an aggregator), and Werner Vogels' "verification debt" remark, which the drafting pass discovered was **misdated by the research pass** — re:Invent December 2024 in the research comment; the primary transcript was not found, but three independent secondary reports converge on **re:Invent 2025**, his final keynote, and this paper uses the corrected date and does not claim the primary. The drafting pass also independently re-verified the METR ~9%-of-task-time figure and the "Coding with 'Enemy'" 56%-still-accept figure directly against source, since the research comment's phrasing of both ("67% minimal code review," "even when a monitor raises a correct alert, sabotage still succeeds in 56% of sessions") could not be reproduced verbatim from the primary on a second read; this paper states only the figures it could itself confirm (56% of participants still accepted flagged code; "minimal code review" as a named but unquantified contributing factor).

**Four evidence kinds**, used throughout: **study** (a controlled or observational measurement with a stated method), **vendor telemetry** (a vendor's measurement of its own users' real behaviour at scale — not controlled, not neutral, but not a marketing claim either), **vendor claim** (a self-report survey or press figure with no behavioural measurement behind it), **author reasoning** (an analytical model or argument, explicitly not a measurement). The ticket's original three kinds could not classify Anthropic's or Faros' telemetry without either overstating it as a controlled study or understating it as marketing; this fourth kind was added on the same reasoning the sibling ladder paper used.

**The nine-candidate Harbour coverage map**, current code checked against HEAD (`b5a4f77f` / `d0e809e3`; cited paths *did* move after this ticket was filed — `git log --since=2026-09-19T10:32:42.717Z` returns four commits touching anchors below, `ec37531d`, `c3d062e6`, `978dddd3`, `95df7ca5`, and `ec37531d` created `docs/ladder.md` v2 outright — but every anchor is pinned at `b5a4f77f`, which post-dates all four commits, and this paper re-read each anchor there, so no verdict below is stale):

| Candidate | Harbour today | Anchor | Verdict |
|---|---|---|---|
| Tests/CI as substitute for reading | The `### What CI Did Not Prove` ledger, with the explicit floor that green CI never discharges a ledger item on its own | `lib/prompt-template-defs.js:1061` | covered — the floor is stricter than the literature's practice |
| Diff size / smaller PRs | No cap, no measurement, no size-triggered lane. Review lanes are "keyed on the risk surface, never on LOC/t-shirt size/file type" | `docs/architecture/prompt-system.md:28` | untouched — and the evidence (finding 2, Kudrjavets et al.) says leave it that way |
| Plan-first / co-planning | A dedicated `Plan-review Gate` with a recorded yes/no decision before implementation | `lib/prompt-template-defs.js:302-313` | covered |
| Sandboxing / permission modes | A single hardcoded launch flag (`--dangerously-skip-permissions`); `permissionMode` is not a stored or varying field anywhere in either repo | `simple-dispatcher/executors.js:483` | untouched |
| Provenance / citations | A mandatory `cites` header and "a citation a reader cannot land on is not one," plus a mechanical gate requiring a landable change-reference before a periodical review can self-conclude | `docs/papers/standard.md:7-21`; `lib/periodical-report-gate.js:20-35` | covered — and this is the candidate with the strongest published evidence behind it |
| Review tooling / trajectory summaries | Review is write-only (ledger + conditional verdict); close-out is a separate step that consumes the ledger and owns the irreversible merge/Done transition | `lib/prompt-template-defs.js:1080-1099` | covered |
| Smaller, bounded tasks | A task-budget guard (`maxTasks`) enforced at the dispatch seam before a bootstrap credential is minted; the bound is persisted per run | `lib/dispatch-factory.js:317`; `lib/dispatch-store.js:379` | covered |
| Graduated / tiered oversight | Not a general tier system; the nearest instance is the Flight Companion's read-only → supervised-writes → unattended progression on one surface | `docs/ladder.md:42` | partly |
| Guardrails, roles, "golden paths" | The two-path prompt system (handwritten + AI-generated, both updated together) is the golden path every dispatch runs through | `docs/architecture/prompt-system.md` (file-level) | covered |

**One deviation from the standard.** `docs/papers/standard.md@b5a4f77f:22-31` fixes a five-part order — Answer, Findings, Method, Limits, Next. This paper has six, with "The strongest disconfirming sources" between Findings and Method, for the same reason and under the same ruling `docs/papers/harbour/developer-adoption-ladder.md@b5a4f77f:55` already cites: John required the six-part form on LIN-2925's ticket (comment, 2026-09-19T10:11Z) rather than let a supportive Findings section absorb the disconfirming material, and the open question of whether every paper should take this shape is already a live line in `proposals.md` (`docs/papers/proposals.md@b5a4f77f:141-146`) that this paper does not re-file.

## Limits

**This paper cannot prove that no study exists showing a verified artifact changes supervision — only that a systematic search across two research passes, dozens of queries, and one dataset checked specifically for this purpose did not find one, and that a live position paper in the field names the same measurement as missing.** A study published after 19 September 2026, or indexed outside the venues searched, would not appear here.

**Three of the cited sources are vendor telemetry and one is an analytical model presented by its authors as such; none is neutral population data.** Anthropic's auto-approve curve — the closest thing to supportive evidence for "verified beats claimed" in this paper — measures Anthropic's own users over a window in which the underlying agent's capability also changed; Faros and Sonar measure their own customers and survey panels respectively, both of which are more AI-engaged than the general developer population by construction; the 84–97% graduated-oversight figure is its authors' own model, not a deployment measurement, and this paper's citation says so every time the figure appears.

**The Harbour coverage map is this paper's own reading of its own code, not checked by an independent second paper.** `standard.md`'s own rule is that a paper is checked by a second paper, not its author; this map should be treated as a first pass until one exists.

**One source is read at one remove by necessity, not by choice.** Werner Vogels' "verification debt" quote is landed via three converging secondary reports (Sonar's own usage, TFiR's reporting, and Implicator.ai's keynote coverage), not an AWS primary transcript; a later edition should land the primary if AWS publishes one. The date correction (re:Invent 2025, not 2024) is this paper's own finding, made by checking the secondary sources' own dating rather than accepting the earlier research comment's figure.

**Dates matter and move fast.** The oldest figures here (METR, arXiv 2606.05391) are over a year old at publication; the newest (arXiv 2609.03460, docs/papers/proposals.md's own live line about this paper) are from the same week this paper was drafted. A reader comparing two numbers across this paper should check the dates before reading a difference as disagreement, the same caveat the sibling ladder paper carries.

**A documented conflict with a sibling paper, not resolved here.** `docs/papers/harbour/developer-adoption-ladder.md@b5a4f77f:73` states "Harbour records permission mode on every run and can compute the same curve for its own users" — a claim this paper's own audit of `simple-dispatcher/executors.js:483` and a repo-wide search for `permissionMode` (zero hits in either repo) finds is not true today: permission mode is a single hardcoded constant with no stored field and no variance to measure a curve against. `docs/ladder.md@b5a4f77f:54` carries the same assumption forward ("the share of runs in each permission mode by session count... computable from what every run already records"). This paper does not amend either document — `standard.md`'s rule is that a paper is revised by rewriting it, and neither file is in this PR's scope — but records the discrepancy here as the concrete instance of "a paper is checked by a second paper" catching a first-paper overclaim, so a future reader of the ladder papers does not inherit it silently.

## Next

**The within-Harbour measurement this paper's ticket asked for is not computable today, on either axis.** Permission mode has zero variance and zero storage (`simple-dispatcher/executors.js:483`); supervision time is not recorded as a concept — the nearest field is `humanContinued`, a boolean per session (`simple-dispatcher/state-store.js:34`), not a duration or a count, and there is no per-operator identity anywhere in the system to attribute it to (`lib/escalation-kpis.js:27`). Proposing to regress supervision time against permission mode as if the data existed would be proposing a measurement with a single-valued independent variable — exactly the failure mode "pin the measurement" is meant to catch. **This paper does not build the field**: adding `permissionMode` is instrumentation for a future measurement, not part of a docs-only paper, and the line below proposes it as the next, separate step.

One line leaves `proposals.md` and one joins it in this PR: the spent LIN-2931 line is removed, and one new line proposes building a `permissionMode` field (stored and forwarded, same shape as `harness`/`effort`) plus a real supervision-time record, before the cost-vs-supervision measurement this paper's ticket asked for can run at all.
