---
title: What does a person expect while one task runs, and what must a checking surface show for an informed merge?
version: 1
date: 2026-09-19
authors: [Claude]
model: claude-sonnet-5, Claude Code CLI, effort default; dispatched as LIN-2946 research item ff48dd04, implementation item 1aa80039
grounded_at: d249ec51 (LinearViewer), d0e809e3 (simple-dispatcher)
cites: [docs/v1.md@d249ec51:47-50, docs/v1.md@d249ec51:54-65, docs/v1.md@d249ec51:99-102, docs/escalation-philosophy.md@d249ec51:50-58, docs/escalation-philosophy.md@d249ec51:115-133, docs/escalation-philosophy.md@d249ec51:134-152, docs/escalation-philosophy.md@d249ec51:153-173, docs/north-star.md@d249ec51:11, lib/prompt-template-defs.js@d249ec51:1061-1099, lib/render-observation.js@d249ec51:1-9, lib/render-observation.js@d249ec51:180-183, lib/render-session.js@d249ec51:88-113, lib/render-session.js@d249ec51:337-397, server.js@d249ec51:1964-1990, lib/kpi-stats.js@d249ec51:1-16, lib/escalation-kpis.js@d249ec51:9-27, lib/task-cost.js@d249ec51:1-25, public/observation.js@d249ec51:1243-1250, simple-dispatcher/feedback.js@d0e809e3:28-31, docs/passage-planner-session-2026-08-03.md@d249ec51:80-95, docs/passage-planner-session-2026-08-03.md@d249ec51:123, docs/passage-planner-session-2026-08-03.md@d249ec51:137, docs/passage-planner-session-2026-08-03.md@d249ec51:250-256, docs/papers/harbour/developer-adoption-ladder.md@d249ec51:13, docs/papers/harbour/what-lowers-the-verification-cost.md@d249ec51 (whole paper), docs/papers/harbour/where-harbour-joins.md@d249ec51:13, docs/papers/harbour/where-harbour-joins.md@d249ec51:68-71, docs/papers/harbour/review-consumption.md@d249ec51:14-18, docs/papers/standard.md@d249ec51:23-31, LIN-2925 (comment, 2026-09-19T10:11Z, ruling `lin2925-paper-order`), arxiv.org/abs/2606.18671 (HANSEL, via what-lowers-the-verification-cost.md), arxiv.org/abs/2606.05647 (via what-lowers-the-verification-cost.md), arxiv.org/abs/2602.16844 (via what-lowers-the-verification-cost.md), arxiv.org/abs/2609.03460 (via what-lowers-the-verification-cost.md), arxiv.org/abs/2203.05045 (Kudrjavets et al., MSR 2022, via what-lowers-the-verification-cost.md), Bacchelli & Bird, "Expectations, Outcomes, and Challenges of Modern Code Review" (ICSE 2013, sback.it/publications/icse2013.pdf), arxiv.org/abs/2605.05564 ("Is this Build Failure Related to my Patch?"), LIN-2946 (research comment, 2026-09-19T20:50:02Z; plan comment, 2026-09-19T20:55:17Z)]
---

# What does a person expect while one task runs, and what must a checking surface show for an informed merge?

No display promises cheaper checking [sourced]: the only candidate with a controlled positive result behind it is provenance — surfacing the evidence for a claim — not tests, not CI, not a smaller diff, and no published study shows that a stronger verification signal changes how much anyone actually supervises. What a run must show is narrower and better evidenced than "more detail": the four moments a person needs to be oriented (what is happening, what comes next, what it has cost, that nothing has been merged), an honest account of when it may ask, and a finished-run page that gives a reader who never built Harbour the same thing a good PR description gives a reviewer — what instigated the change, what was checked, and what they are trusting when they click. Harbour already has pieces of every one of these; what it does not have anywhere is one page scoped to one task, showing cost and merge state, readable by someone signed out. Which of the three candidate shapes assembles those pieces, and how, is not this paper's question — it hands that to the sketches sitting (LIN-2947) with the evidence and the open questions below.

## Findings

**1. The three people papers license a narrow claim, not a broad one, and one of them leaves a question open. [sourced]** `developer-adoption-ladder.md` found the population stalls at handing over without watching, not at using an agent at all — every source that names a cause names the cost of verifying [developer-adoption-ladder.md:13]. `what-lowers-the-verification-cost.md` found the only candidate for lowering that cost with a controlled result behind it is provenance (HANSEL, arXiv 2606.18671, n=14, web-browsing agents) — not tests, not CI, and smaller diffs has a *negative* result against it (Kudrjavets et al., MSR 2022, 845,316 pull requests, no measurable relationship between PR size and time-to-merge). `where-harbour-joins.md` found Observation cannot be placed on the developer-adoption ladder without a ruling from John: ladder v3 gives it no rung at all, while John's own framing is that Harbour "first steps in as observation" [where-harbour-joins.md:13, :68-71]. **This paper does not resolve that disagreement** — it is `lin2930-observation-rung`, an open question this paper hands forward unchanged, alongside this paper's own closing agenda for LIN-2947.

**2. Of the four things a person must know at every moment, today's surfaces serve one in part and the other three not at all. [John, 19 Sep 2026, `docs/v1.md:47`, for the four needs; sourced for the "what exists today" column]** John's own words, quoted in step 5 of the person's path: *"Just sees a lovely interface of the work in progress."* At every moment the person knows what is happening, what comes next, what it has cost, and that nothing has been merged [`docs/v1.md:47`].

| Need (`docs/v1.md:47`) | What exists today | Anchor | Verdict |
|---|---|---|---|
| What is happening | Session feed: status pill, one-sentence summary, runtime + model, per-run progress; drill-down to phase/recap/metric chips and the activity log | `lib/render-observation.js:1-27` | Partly — present, but keyed on the session spine, not on one task |
| What comes next | Nothing names the next step to a reader. Per-run cards carry `kind`, iteration number, waiting/parked flags, and an agent-to-agent handoff marker | `lib/render-session.js:337-397` | Absent for a reader — inferable only by someone who already knows the pipeline |
| What it has cost | Not rendered on either watching surface. The data exists as `kind:'usage'` on the feedback stream and as a joined per-task USD figure; the only rendered money anywhere is the instance-wide KPI card | `simple-dispatcher/feedback.js:28-31`; `lib/task-cost.js:1-25` | Absent from the run view — a display gap, not a data gap |
| Nothing has been merged | No merge-state indicator anywhere on either surface. Produced-artifact URLs matching a `/pull/N`-shaped path are classified as a PR and linked out; merge state itself lives on GitHub, not Harbour | `public/observation.js:1243-1250` | Absent |

**3. When a run may ask is already answered and ratified — this paper cites it rather than re-deriving it. [sourced]** `docs/escalation-philosophy.md` defines an escalation as anything requiring a human response (Principle 0), requires de-duplication to root cause [escalation-philosophy.md:115-133], names three nuisance patterns to kill — chattering, standing/stale, fleeting [escalation-philosophy.md:134-152] — and requires every escalation to be self-sufficient: what, why, the decision, the options, and the cost of each path including doing nothing [escalation-philosophy.md:50-58, :153-173]. False-escalation rate — dismissed over answered-plus-dismissed — is already computed [`lib/escalation-kpis.js:9-27`] and is a headline KPI at the normative layer: *"False-escalation rate is a headline KPI: every surface that asks a human is judged by how often the answer was 'why was I asked this?'"* [`docs/north-star.md:11`]. **[John, 19 Sep 2026, `docs/v1.md:60`]** states the same requirement in the person's own terms: the person is never asked something they cannot answer, and never asked twice; false-escalation rate is a headline measure. Today, the mechanism that answers this is a tab in the **operator's** Observation view [`lib/render-observation.js:180-183`] — the doctrine is right, the inbox is wrong for a user's own run.

**4. The finished run's first screen, for a reader with no repo access, is closer to a PR description than to a diff or a transcript. [sourced]** Bacchelli & Bird (ICSE 2013, Microsoft; observations, interviews, surveys, hundreds of classified review comments) found that change understanding is the key work of review, and that the single biggest reported information need is "what instigated the change" — the tool they studied puts an author-written `description.txt` first, before the file list. Harbour's own thing a reader most needs — the `### What CI Did Not Prove` ledger, with each item marked inside or outside the ticket's scope and a named discharge route, and a verdict that must read `Approve — conditional on close-out discharging the ledger` whenever the ledger is non-empty — has **no rendered surface today**; it is prose written into a tracker comment [`lib/prompt-template-defs.js:1061-1099`]. [author] Any display of it must preserve those two marks and the conditional-versus-plain distinction; a checklist that drops the discharge route or reads a conditional Approve as unconditional has flattened the thing a reader was trying to check.

**5. The reader this paper is about — someone with no repo access — cannot reach today's run pages at all, and the share link is a new boundary, not an extension of one. [sourced]** `server.js`'s `workspaceFromUrl` middleware redirects any workspace page to `/`, and 404s an unrecognized key, whenever the session carries no connected workspace [`server.js:1964-1990`]. The only public, unauthenticated read surface today is `/kpis`, whose collector states its own scope as the privacy boundary: "only counts, day buckets, and app-defined labels ... never workspace urlKeys, prompt text, summaries, tokens, or issue content" [`lib/kpi-stats.js:1-16`]. A run page that a share-link holder can read is not a wider `/kpis` — it necessarily shows issue content, which `/kpis` is built to forbid. [author] The paper treats the signed-out share link as the **first deliberate crossing** of that boundary, decided once by whoever designs it, not as `/kpis` already having done the work.

**6. Observation and the session pages are built for, and speak the vocabulary of, the operator — not the reader this paper is about. [sourced]** Four tabs — Autopilot, Sessions, Rulings, Scan-due [`lib/render-observation.js:180-183`] — are four pieces of Harbour's own vocabulary before any content loads. Session-run chips read `⏱ runtime`, `◐ N heartbeats`, `✎ N artifacts`, `◇ model`, `▤ N MB peak`, `⚿ credential` [`lib/render-session.js:88-113`]; peak memory and credential state are operator telemetry, and a heartbeat count is a liveness concept with no meaning to a first-time reader. The unit both pages organize around is the **session** — a reconstructed `sessionId` spine spanning a seed task's descended and spun-off children, explicitly not one task [`lib/render-observation.js:1-9`] — while `docs/v1.md`'s run is exactly one task. [author] "Watch the run" and "the PR with its evidence" both already appear as named gaps in `docs/v1.md`'s own step table: "The Observation view and session pages, built for the operator" against "The run experience, designed for a person who did not build Harbour" [`docs/v1.md:99-102`].

**7. The three candidate shapes, scored against findings 2–6, without choosing one. [author]** `docs/v1.md:65` names three candidates and says the design track starts from the person's expectations, not from any of them:

| Candidate | What it already serves (findings above) | What it still has to build |
|---|---|---|
| A conversation at task scale, evidence landing in-thread | Finding 3 (an escalation can carry its own why/options/cost inline); this shape already exists once, in Flight Companion's live Approve/Dismiss proposal chat, and once more, task-scoped, in Task Chat | Findings 2 (cost, merge state), 5 (signed-out read), 4 (a first screen that leads with intent, not the latest turn) |
| Expected steps as things to press or let run | Finding 2's "what comes next," made literal as a control | Finding 4 (a finished-run summary distinct from a step list); risks reproducing defect 1 below if pressing becomes a reflex |
| One page per run, CI-pipeline-shaped | Finding 4 (a page-first, not thread-first, summary — closest to the Bacchelli & Bird "description first" finding); nearest to the outside surfaces in §Method | Finding 3 (a pipeline page has no natural place for a mid-run question); Finding 6's own defect risk (CI pages are read for status, not for what instigated the change — see the disconfirming case) |

No shape is recommended here. Each serves some of findings 2–6 and costs something against the others; which one John and the sketches sitting choose is `docs/v1.md:65`'s call, not this paper's.

## The disconfirming case

**Harbour has already run something close to the first candidate shape once, live, and it produced the merge click's own failure mode one milestone early. [sourced]** The first live passage-planner session (2026-08-03) recorded seven human-interface defects, all in the interface layer, none in the evidence machinery [`docs/passage-planner-session-2026-08-03.md:80-95`]. The sharpest: serial, one-leg-at-a-time ratification with no whole-plan context produced rubber-stamping — John, contemporaneously, *"honestly I'm just clicking approve without reading"* — alongside orientation that happened but was never shown before he was asked to ratify, and a false ratification record claiming "all four legs are ratified" off hollow per-leg clicks, with no way to tell a genuine yes from a fatigued one. Session #2, after the interface was revised (orientation shown first, the whole proposal shown at once, evidence gaps stated up front), read as *"much stronger"* [`docs/passage-planner-session-2026-08-03.md:123`] — but the **challenge/negotiate half of the fix was never exercised** in either recorded session [`docs/passage-planner-session-2026-08-03.md:137, :250-256`], so half of the mechanism that is supposed to prevent rubber-stamping ships unvalidated. [author] Defect 1 is a live in-repo instance of the same failure this paper's outside sources describe below: an approval surface that asks often enough, with little enough context, converts a decision into a reflex.

**"Assign, walk away, review the PR" is the live hypothesis that the watch surface, whatever shape it takes, may never be opened at all.** [author, drawing the inference; this paper cites no source for how common the pattern is] If that is how people actually use Harbour once trust is established, everything rides on the finished-run page (finding 4), and effort spent making the in-progress watch view richer would not be effort spent where the checking actually happens.

**Two controlled studies that manipulated a verification signal found behaviour barely moved, and a third found the signal moved the wrong direction. [sourced]** "Coding with 'Enemy'" (arXiv 2606.05647, 100+ developers): 94% missed deliberately planted agent sabotage unaided, and even after an automated monitor correctly raised an alert, **56% of participants still accepted the flagged, malicious code**. Microsoft Research's three-study series (arXiv 2602.16844, n=12 per study): a step-focused trajectory view "can justify overreliance," and a better interface made people find errors faster and report more confidence — with **no meaningful effect on accuracy**. arXiv 2609.03460 (n=81) measured a "transparency penalty": disclosure that content is AI-generated lowers perceived trustworthiness at constant quality; only a much denser, visualized provenance signal restored real discrimination between accurate and fabricated content, and a thin signal (a disclosure label alone) moved trust the wrong way.

**A reader of Harbour's own evidence today consumes the ledger almost completely and the narration around it barely at all — the one measured precedent says more detail is not more checking. [sourced]** `review-consumption.md` found that across ten final reviews, a close-out reads 97% of the `### What CI Did Not Prove` ledger's sentences but only 25% of the method and check narration, the review's largest section [`review-consumption.md:14-18`]. The reader in that study is a machine, not a person — but it is the only reader of a Harbour evidence artifact Harbour has ever measured, and its reading pattern argues for a page that leads with the ledger and the verdict, not one that narrates the run in order to look thorough.

**This paper's own central claim is disconfirming of the instinct to build a richer display.** No published study shows that a stronger verification signal changes how much anyone supervises [`what-lowers-the-verification-cost.md`, whole paper]. A run page that adds detail on the theory that detail lowers checking cost is asserting something three cited studies found unsupported or false.

## Method

**Population and classes.** This paper does not run its own study; it synthesizes three existing Harbour people-papers (`developer-adoption-ladder.md`, `what-lowers-the-verification-cost.md`, `where-harbour-joins.md`), reads two outside sources in full this session (Bacchelli & Bird, ICSE 2013; the CI-attribution paper below), and reads Harbour's own code and doc anchors at HEAD. Four bounded classes, inherited from the research and plan and re-verified here: **reader roles** (5 — operator, the v1 user who pressed Go, the share-link holder, machine readers, the invited tester — bounded from `docs/v1.md` and its milestone children; not boundable from the repo is any real user outside John, since none exists yet); **run-reading surfaces** (bounded by enumerating `lib/render-*.js` at HEAD; this paper's Findings draw on Observation and session directly and name Flight Companion, Task Chat, and Live Console once each as already-existing pieces of the candidate shapes, without individually re-analyzing the rest); **evidence and provenance representations** (the ledger, the verdict, the close-out summary, the feedback stream's `evidence` kind, produced-artifact links, cost); **moments a run can ask** (5, enumerated in the LIN-2946 research comment of 2026-09-19T20:50:02Z; each tested against Principle 0 at `docs/escalation-philosophy.md:50-58`, which is the test for what counts as an escalation, not the enumeration — the doctrine is cited, not re-derived).

**Outside sources read.** Bacchelli & Bird, "Expectations, Outcomes, and Challenges of Modern Code Review" (ICSE 2013, Microsoft; observations, interviews, surveys, hundreds of classified comments) [study]. "Is this Build Failure Related to my Patch?" (arXiv 2605.05564, 77,354 CI build failures across seven projects, 371 manually analysed): developers spend a median of four hours establishing whether a failure is even related to their own change [study]. [author] The reading cost on a CI page is attribution, not status — a red step that does not say whose problem it is has reproduced this paper's own finding about the ledger's discharge route being the thing that must not be dropped.

**A leg this paper deliberately narrows.** The research pass for this paper searched for how people read deploy logs and incident postmortems and found mostly practitioner blogs with no measured reading order behind them. That material is not cited here: it did not clear the bar this paper otherwise holds, and a claim about deploy-log reading order would be [author, on method] at best, not [sourced]. This paper omits the leg rather than let a plausible ordering pass as evidence.

**Source marking.** Every sentence in this paper asserting what a person wants or expects carries one of three marks: **[sourced]** — a study, an in-repo dated record, or a measured Harbour finding; **[John, dated]** — John's own stated view, quoted; **[author]** — this paper's own reading or inference, never itself evidence. The marks are inherited from the LIN-2946 research comment (2026-09-19T20:50:02Z) rather than re-derived, per the standing convention set by `what-lowers-the-verification-cost.md`'s four evidence kinds.

## Limits

**One human reader.** Every in-repo record of a person reacting to a Harbour-built approval surface is John, on one occasion (the passage-planner sessions of 2026-08-03). No user other than John has ever read a finished Harbour run. Every claim in this paper about what "a person" wants rests on John plus the published population sources — not on any Harbour user's own account.

**No transcript exists for either passage-planner session**, despite the ticket that commissioned it calling for one; the disconfirming case's account of session #1 and #2 rests on a contemporaneous chronicle written from notes, not a replayable record [`docs/passage-planner-session-2026-08-03.md:250-256`].

**The challenge/negotiate mechanism remains unvalidated.** It shipped in the interface revision that produced the "much stronger" verdict, but was never exercised in either recorded session — half of the fix for rubber-stamping is untested in production, and this paper cannot say whether it works.

**The deploy-log leg is omitted, not disconfirmed.** This paper found no controlled or observational study of incident-reading order strong enough to cite; a second pass that lands a primary source (SRE/incident-response literature, or an observability vendor's own telemetry) could still add this leg with proper standing.

**`lin2930-observation-rung` is left open on purpose.** `where-harbour-joins.md` states the disagreement between John's framing and ladder v3's silence and hands the ruling to John; this paper adds nothing to that disagreement and does not adopt either reading.

**No instrument in Harbour measures whether checking got cheaper.** `permissionMode` has zero variance and zero storage; the nearest field to a supervision measure is `humanContinued`, a boolean per session, not a duration; there is no per-operator identity separate from per-workspace. The one live, honest instrument is false-escalation rate, and it measures the asking side of the run experience, not the checking side.

## Next

Apply `review-consumption.md`'s method to a human reader: on the first invited runs of a finished-run page (whichever shape the sketches choose), record which sections a real reader opens, how long they stay, and whether they reach the ledger before clicking merge, then report the share nothing reads — the same question `review-consumption.md` answered for a machine reader, asked for the first time of a person. (Claude, 2026-09-19)

---

## Questions for the sketches conversation (LIN-2947)

*This closing agenda is a section this paper carries after `Next` because LIN-2946's own Shape section requires it. It is a one-off for this paper, not a change to the paper standard: `docs/papers/standard.md` and the standing ruling `lin2925-paper-order` still fix the order at six parts ending in `Next`. Whether the standard should sanction an agenda section at all is John's call, routed to the LIN-2947 sitting and open until he takes it.*

1. **One page or one thread?** Is the run a conversation that accumulates evidence in-thread, a pipeline page that updates in place, or a page with a conversation inside it?
2. **What is the unit?** `docs/v1.md`'s run is one task, but today's Observation/session spine is the session (Finding 6). Does the run view get its own unit, or does the session spine get re-scoped?
3. **What does the first screen show before anything is pressed**, given that the planner's worst recorded defect was orientation that happened but was never shown (disconfirming case)?
4. **What may the run ask, and where does the answer land?** Which of the five doctrine-bounded moments is permitted on a user's own run, and does the question reach the user directly or the operator's Rulings tab (Finding 3)?
5. **What is the one number on the screen while it runs** — money spent, budget remaining, steps done of expected, or elapsed time? The data exists as `kind:'usage'`; the display does not (Finding 2).
6. **How does "nothing has been merged" get said**, and does the page reflect live repository state or only its own record of what it did?
7. **What is the finished run's first screen for someone with no repo access** — the intent-and-outcome statement, the ledger, or the diff (Finding 4)?
8. **How is the ledger displayed without flattening it?** The inside/outside-scope mark, the discharge route, and the conditional-versus-plain Approve distinction all have to survive (Finding 4).
9. **Where does the merge click live, and what must be true on the page before it is pressable** — the anti-rubber-stamping question, with defect 1 and the 56%-still-accept finding on the table (disconfirming case).
10. **What does the share link show, and what does it withhold?** `/kpis`'s boundary forbids issue content; a readable run page cannot honour that boundary unchanged (Finding 5) — what is the deliberate new one, and who can turn it off?
11. **Do the hands-on person and the handover person really share one screen**, or is the prompt path (`docs/v1.md`'s step 4 aside) a second door on the same page?
12. **The standing ruling:** which rung, if any, observation belongs to — `lin2930-observation-rung`, already routed to this sitting and left open by Finding 1.
