Do the supervisor, test-estate and throughput papers hold up?
On their headlines, yes; on several load-bearing claims, no. Every committed script re-runs, and
nearly every printed number matches exactly. Supervisors: a fresh blind recode puts the
mechanical share higher, at 86% of tokens (81–91%) against the paper's 77%, so "mostly
bookkeeping" holds. But it holds because the codebook calls a gate reply and a fetch mechanical,
not because of the coders. The quiet-wake share is 30.7% of wakes and 25% of the bill, not 32.5%
and 30%: the paper counted a resume handshake as a wake. "The answers are already held in code"
cites the right lines. But two of its claims are wrong: that the prompts ask for polling
alongside wakes, and that the merge-and-Done gate is only prose. Its list of supervisor slips
does not survive a reading of the tickets, though the total does: 22 of 25 failures were
mechanical, 21 in code. Test estate: the 25-second idle is real and is 31% of serial time,
but armKeepalive does not cause it. A timeout timer in routes/proxy.js that is never cleared
does. Flakes are the largest cause of red CI. The pin evidence is thinner than the paper says:
"pin edits 2.8%, rising" counts whole files, and assertion-literal edits fell. Only about five of
the 18 "bump" tickets are bumps. The catch-against-bump table rests on one or two real catches
a class. Throughput: the printed numbers reproduce, but the definitions carry the
reliability paper's finder-row error. 35 of the 66 Bugs that "name their introducer" name the
ticket whose review found the fault, and the true median lag is 3.3 days, not the same day.
June's 95 a week covers the weeks of 8 to 29 June; June's calendar weeks give 69. And the
per-change detection power treats changes as independent: over four weeks it detects ×2.8 in
dispatches and ×2.0 in hours, not ×1.4 and ×1.3. A doubling still shows first in hours per
change, but only just, and not within one or two weeks. The scorecard's weekly dispatches are
clear of the July run-log gap; the per-ticket July figures are not. Each paper's version 2
corrects its statements; judgement disagreements stay here. Nineteen lines of
docs/steady-base.md change, and a twentieth misquotes the earlier check.
Findings
What supervisors do
Every script re-runs byte for byte, and every printed number matches. A fresh extract over
the same window is identical to the author's cache under cmp: no transcript in scope had aged
out yet. Classify, sample, failures (offline), analyse and render then reproduce every output
file and both committed figures. Over a copy of the author's cache the same holds.
| Paper's figure | Re-run | Verdict |
|---|---|---|
| 32,525 steps, 254 sessions; 2 / 35 / 118 / 96 / 3 by layer | same | confirmed |
| 77% of supervision tokens mechanical, 80% by the second estimate; 61% of working time | 77.1, 80.3, 61.4 | confirmed; the recode puts it higher, below |
| The class table: steps, token shares, mechanical shares, time shares | same, row for row | confirmed as the rules' shares; the recode moves restate and poll, below |
| Median gate reply 34k units against 30k for a beat; 24k → 71k over 56 sessions | same | confirmed |
| 82 hours at work, 1,459 held | 81.6, 1,459 | confirmed |
| A third of wakes quiet, 1,782 of 5,485, 30% of the bill | same | 30.7% and 25% once a handshake is not counted as a wake, below |
| The Runner "re-arms on 94% of its 970 wakes" | 94.1% of its wakes are quiet; it re-arms in 98% | label corrected |
| Stepper wakes 9% quiet | same | 7% without the handshakes |
| Runner 15.8% of the bill, 97% mechanical; leg and stepper 21.6% and 34.4%, 70–71%; autopilot 26.7%, 92% | same | confirmed; these come from the direct estimate, which totals 80%, not 77% |
| Coder agreement 93.6% (κ 0.92), M or J 94.5% (κ 0.83), rules 71.8% (κ 0.67) | same | confirmed |
| 26 failures of 82 candidates; 12 / 6 / 6 / 2 by type; 24 mechanical; 21 in code | same counts | LIN-2145 is not among the 82; the classification changes, below |
| 108 failsafe re-confirms, 40 nudges, 30 refused duplicates, 47 other dispatch errors | same counts | 27 duplicates; the 47 are mostly outages, below |
| 27% of the fleet's weighted tokens | 35 × 0.771 | arithmetic confirmed; the 35% is another paper's, over a wider window |
A fresh blind recode puts the mechanical share higher: 86%, not 77%. 160 steps, none of them among the paper's 220, were drawn systematically with probability proportional to their weighted tokens. On that design the share of sampled steps coded M estimates the token-weighted mechanical share directly, with no rule classes and no layer weights. Two coders coded every card blind to the paper, its codes, the rules and each other, using the paper's own codebook unchanged.
- Agreement. Class 98.1% (κ 0.98), M or J 92.5% (κ 0.68). The rules against the coders, where the coders agreed: 66.9% (κ 0.61), against the paper's 71.8%.
- The share. 86.2% of tokens mechanical, split verdicts counting half; a 20,000-resample bootstrap gives 81–91%. 82.5% of steps were M to both coders. By layer: Runner 100% (n=25), leg 86% (35), stepper 80% (55), autopilot 88% (43).
- The sample matches the census. The rules' classes of the 160 steps are 25% read, 23% re-arm, 20% restate, 12% judge and 7% dispatch, within sampling error of the paper's token shares (23.7, 24.2, 15.4, 13.7, 8.4).
- Where it differs from the paper. By the coders' class, judging a report is 68% mechanical (paper 50%) and writing a beat 45% (paper 22%). Most mechanical judge steps check a PR, CI or status the worker reported. The rules' largest error in this sample is not re-arm but restate: 18 of the 32 steps the rules call restate are acknowledgements of a progress wake or a resume handshake, which both coders call poll. Poll is about 19% of tokens by the coders against 4% by the rules, restate 11% against 20%. Both are mechanical, so the total is unaffected.
- The coders flagged two failures: the 5 September wake doing a worker's job, which the paper
also found, and an autopilot resolving a child's escalation itself instead of putting it to John
(
1de62534:484).
The share is robust to the coders and hinges on the codebook. Varying the estimator moves it little: split verdicts as J 74.6%, as M 79.5%; either coder alone 77.3% and 76.9%; step- rather than token-weighted 78.3%; a within-layer bootstrap 72–82%; without the two Runner sessions 73.6%. What moves it is the definition. The codebook makes a gate reply and a fetch of a known URL mechanical by rule. Without the re-arm class the rest of the bill is 70% mechanical; without re-arm and wake deliveries, 62%. And "read" is not all idle: about 16 of its 23.7 points are reads in a cycle that goes on to judge, dispatch, relay or decide. The finding is exact for what it defines, that three-quarters or more of the tokens pay for steps whose action state decided. Whether code could take those steps over without losing what the model reads along the way is the paper's first Limit, and this check does not settle it either.
The quiet-wake share counts a handshake as a wake. The cycle detector
(scripts/survey-supervise-classify.mjs:27, -analyse.mjs:61) starts a cycle at "resumed to
handle a follow-up … reply 'ready'", which the runner sends just before the "Your task … is
ready" wake it announces. That makes 142 one-step quiet cycles, and because a resumed session
rewrites its whole context cache they cost 5.9% of the bill. Merged into their wakes: 1,639 of
5,342 cycles quiet, 30.7%, and 25.3% of the bill; the Runner 94.0% of 955; steppers 7%.
Leaving orient out of the quiet set changes nothing that matters (32.3%); leaving restate out
too drops the share to 18.8%, because 749 quiet cycles contain a restate step.
- "These are the progress wakes" was asserted, not measured. Reading the wake type from each fetched prompt: 87% of quiet cycles are "paused (pending), not done" progress wakes, and the rest terminal wakes (73), handshakes and task notifications. 39% of progress wakes (980 of 2,525) did lead to an action, so a progress wake and a quiet wake are not the same thing.
"The answers are already held in code": the lines are right, several claims are not. Every cited line was read at its sha.
| Citation | What the paper says | Verdict |
|---|---|---|
hook.js:418, :491, :1062 |
the gate asks, parses, the wake text | correct |
hook.js:1437 |
parks PENDING-EXTERNAL in AWAITING_EXTERNAL |
correct, after a warm hold expires (:1452) |
reapers.js:877, config.js:112 |
fails a session that re-declares PENDING too often | correct (three refires) |
reapers.js:1207 |
failsafe after 60 minutes | correct; the 60 minutes is config.js:50 |
reapers.js:1009 |
live-child knowledge used "only" to exempt from the stall reaper | too narrow: it also gates the expired-hold fail path and the inactivity reap (:2169, :2905). "Never tells the parent" holds |
phases.js:131 |
isLegalTransition has no run-time caller |
correct in both repos |
lib/dispatch-wake.js:75, :151 |
the progress text; guards self-wakes, loops, aborts | correct |
lib/dispatch-factory.js:62, :187 |
refuses duplicates, over-budget and closed-ticket dispatches | overstated: duplicates only within five minutes and without followUpTo or force (:272); budget only when declared (:389-393); closed tickets only on the composed-run path (:513-567). Lineage inheritance is right |
lib/openrouter.js:1799, lib/recommend-recurse.js:106 |
recommend is a model call in deterministic wrapping | correct |
autopilot-kickoff.js:198, :386, :467 |
the child-set; the 30-minute clock; the plan bound in prose only | correct; the clock is at :392-393, the bound at :469-472 |
autopilot-kickoff.js:447; passage-runner-prompt.md:132, :139 |
prompts ask for polling alongside push wakes | wrong: the kickoff forbids polling a subscribed child (:136, :384); :447 lists a watch verb; the passage prompt polls only a leg with no dispatch row, instead of a wake |
passage-runner-prompt.md:56 |
re-read state rather than trust notes, "and keep those notes anyway" | the first half is right, about the passage description; the second has no source |
| "the merge-and-Done ledger gate" is prose only | partly wrong: GitHub refuses a merge without green CI, and the proxy refuses Done on periodical tasks (routes/proxy-writes.js:309-316). The Approve-and-ledger half is prose |
The lede's "nearly all the answers are already held in code" is a claim the paper's own first Next proposes to measure. The code check supports "much of it" for the completion gate, wake delivery and stall handling, and does not support it for the factory's refusals or the polling. Of the 24% "spent re-arming", the part that restates live-child state is the PENDING-EXTERNAL replies, 20.4% of the bill.
The failure list mostly survives in total and not in its examples. Every one of the 26 and a systematic 15 of the 57 excluded candidates were re-read over the proxy with their comments, plus eight targeted tickets. Version 1 classed from the title and the first 1,500 characters.
- Six are not supervisor failures: LIN-1884 and LIN-2899 are worker sessions, LIN-1591 is a measurement (its own correction reads 14% of dollar spend over six sessions), LIN-2754 is a feature gap, and LIN-2229 and LIN-2468 are latent, with no incident.
- Five belong in: LIN-430 (an autopilot merged two red-CI PRs, June), LIN-353, LIN-870, LIN-1280 and LIN-1451.
- Four change locus. LIN-1323's fix was in
hook.js, which returned before parsing a correct PENDING-EXTERNAL; LIN-1656 was a missing creation guard. Both were "model", both are code. LIN-881, labelled code, is a supervisor that did not setsubscribewhen told to. LIN-384 blames the recommend engine's defer decision, a model call. - On that reading, 25: 15 missed or lost wakes, 4 loops, 4 wrong routings, 2 lost inputs; 22 in mechanical actions; 21 in code, 2 in the recommend engine, 2 a supervisor's own model slip. None of version 1's five "supervisor's own mechanical slips" holds up as one. The headline, mechanical and mostly plumbing, stands. Merging LIN-2517 into LIN-2511, which it splits from, gives 24.
- Of the transcript signals, 3 of the 30 "refused duplicate dispatches" are reads of a guide
that mentions
DUPLICATE_DISPATCH; of the 47 "other dispatch errors", 34 are HTTP 503s from tracker authentication or the AI service and 8 are failed reads.
Smaller slips, each corrected in version 2. There are twelve classes, not eleven. The relay
example is backwards: both sampled relays of a ruling were J to both coders, and the one M relay
points a gate line at an open decision. The mechanical-beat example ("sends preplanned beat 4/4")
was a split verdict; the three beats both coders called M re-send the engine's recommendation.
review-loops.md counted 83 of 94 tickets sent back at least once, not plans. Not every number
is printed by the analyse script. Two cited papers are not used and nine named tickets are not
cited.
The test estate
The measurements re-run; several of the readings built on them do not. The shape script was
re-run at HEAD (no test changed since c65b7dd8), the timing script as a fresh serial run in a
scratch worktree at c65b7dd8, CI fetched and classified fresh through gh, the session script
fresh, the friction scripts at c65b7dd8, and the analysis over a copy of the author's cache.
Seven Harbour logic mutants (every fifth) and all four killed prose mutants were re-run.
| Paper's figure | Re-run | Verdict |
|---|---|---|
| Class table; 14,097 cases, 12,451 unit; 851 pins; 93% behavioural | identical per test | confirmed |
| Blind validation 64 of 85; pins 22 of 25; behavioural 35 of 36 | identical (--validate) |
confirmed |
| Serial 324 s, 74 s inside tests (23%), median file 0.28 s | 323 s, 72 s (22%), 0.277 s | confirmed |
| Four files idle ~25 s, 101 s, 31% of serial | 101.7 s, 31.5%, the same four files | confirmed; the cause is wrong, below |
| Pin classes 18 s, about 5% | 17.3 s, 5.4% | confirmed |
| 181 failed attempts in 3,190 LinearViewer runs | 183 in 3,216 | the snapshot missed 26 runs of 12 and 15 June |
| PR attempts 54 flaky / 34 test / 19 fault / 5 bump / 10 unresolved | same from the cache | confirmed; recounted by run, below |
| Sessions 1,219 / 11,088 runs / 6,169 episodes / 3,267 deliberate / 1,386 test-first / 225 | identical up to the author's cut | confirmed as counts; the labels, below |
| Friction: 2.8%, 1.0 → 4.9%; 32 PRs; 18 tickets, 10 in September; 61 | identical | confirmed as counts; what they measure, below |
| Mutants 28 of 32; prose 4 of 19; simple-dispatcher 37 of 43 | identical; the 11 re-run match outcome and killers | confirmed |
| Fault label right 10 of 11 on PRs, 2 of 9 on main | no committed source | cannot be re-run |
The idle is a never-cleared timeout, not armKeepalive. The four files
(lin-2216-transient-vs-terminal-auth, lin-2804-provider-ui-endpoint-hints,
proxy-openrouter-fallback-note, lin-2363-kickoff-provider-attribution) each take 25.5 s alone.
They take 0.4–0.5 s with --test-force-exit, and their longest test is 0.1 s. None imports
armKeepalive; all import routes/proxy.js. A preload that traced setTimeout found each file
arming a 25,000 ms timer in withTimeout (routes/proxy.js:374-384, set from
GRAPHQL_TIMEOUT_MS at :337), reached from resolvePromptIssueContext, and never cleared. In
a scratch worktree, unref'ing armKeepalive's timer left all four at 25.4–25.6 s. Unref'ing
withTimeout's brought them to 0.43–0.45 s with every test passing. The two 25-second defaults
are a coincidence. The paper speaks of serial time, correctly. It does not say the idle costs
wall-clock time too. A greedy pool over the per-file times finishes in 47 s with it and 26 s
without it at nine workers, and in 108 s against 75 s at three. CI runs the unit suite twice.
Measured whole-suite runs on the loaded runner machine pointed the same way, noisily. LIN-3158,
merged while this check ran, clears the withTimeout timer on settle.
Flakes are the largest cause of red CI, by attempt and by run. The classifier calls an
attempt flaky when the next green run has the same sha. Every one of the 72 flaky rows is a
re-run of the same workflow run, so it tested identical code. Counting attempts inflates flakes: one run re-tried to
attempt 6 gives five rows. By run, 176 failing runs: 47 e2e flaky, 28 e2e fault, 28 e2e test
updated. Faults and test updates together still outnumber flakes. A hand check of every 9th
LinearViewer PR attempt (15; 3 logs expired) agreed with the label 13 times. One flake had been
labelled a merge from main. The same observation-rulings test failed on two branches on the
same day, which suggests a real race rather than a harness flake.
Several CI statements are wrong in detail.
- "Unit-suite reds are mostly flakes": 13 of the 42 unit-job attempts are flaky and 13 are faults.
- "14 attempts fixed in interface code" are fewer faults. Three
header-navattempts are one branch failing a test its own PR added, twonext-runattempts share a commit, and two were fixed in backend code. - "Failing tests by class, e2e 107, behavioural 31" is both repos' 140 PR attempts, including 38 with no named test, counted by job. LinearViewer's named tests alone give 71 and 24.
- "Pins turned CI red on a PR only twice" is true of the 851 unit pin-class tests. Count pins
inside browser specs failed PR CI three more times (
audit.spec's template count twice,templates.speconce). They are classed e2e by the first matching rule.
The pin evidence inside sessions rests on a handful of events. The per-thousand table turns 4, 4 and 12 text-pin episodes and 3, 5 and 5 census episodes into rates. All 33 were read by hand:
- Census: one real catch against three bumps. None of the author's five "bumps" is a bump (two are a test-authoring error, three a missed probe), and two "rewrites" and one "fix" are bumps.
- Text pins: two catches, one a syntax error, one bump, and three false positives from an ordering pin. All 12 "rewrites" are one mutation probe the detector missed. So "rewritten three times as often as either" is wrong.
- Direction: "about as likely to demand a bump as to catch a fault" holds, bumps at or above catches, on one or two catches a class.
- Source scans: 6 of their 12 "fixes" are one session re-pointing scans after code moved, the caveat the paper itself raises.
- Behavioural: "they make almost every catch that survives" reads the 225 production edits, which the paper rightly calls a ceiling, as catches. Every 16th behavioural episode was read (10): one was a real catch, the rest missed probes, tests written first, the AppleScript environment failure, a date-dependent test and a fixture.
- The session window is 31 August to 30 September, not 30 July: 30 July is one stray run.
The friction figures measure less than they say.
- "Pin edits are 2.8% of changed test lines, rising to 4.9%": 88% of those lines are every changed line of a file flagged as a pin file. A file is flagged for as little as the word "census" or one source read. Assertion-literal edits alone are 0.33%, and they fell from 0.75% in June to 0.27% in September.
- The 18 tickets, each re-read, are mostly not bumps. About five are: two snapshot re-baselines (LIN-1033, LIN-1614), two doc or comment counts with no test (LIN-687, LIN-2481), and LIN-3133. The other 13 repair or extend weak pins. Of the "four pure count bumps", LIN-2985 and LIN-2989 add entries nobody had pinned, LIN-2988 was cancelled, and LIN-3133 also changed production code. The paper's own rule misses two more (LIN-2554, LIN-2590): 20 in all, 12 in September. One of the 18, LIN-2903, is simple-dispatcher's, against "no ticket".
Smaller slips, each corrected in version 2.
opencode-runner.test.jsfails whenSD_WORKER_USAGE_RELAY=1is set, notHARBOUR_LOCAL_BASE.- Behavioural tests caught two of the four killed prose mutants, not one.
- Two of simple-dispatcher's 47 curated mutants target Harbour and were not run: 45 ran, 43 still apply.
- The 6,234 "text pins" are
growth-atlas.md's count, notsurvey-check.md's. - Two citations are a few lines off (
cheap-implementer.md:208-210,growth-atlas.md:104), andreliability-baseline.md:106is cited but not used.
Measuring throughput
Every printed number reproduces. The git, runner, fleet and token scripts were re-run fresh
with origin/main pinned to the paper's shas. The scorecard was run over the author's tracker
and GitHub snapshots, and 19 fresh proxy reads checked the ticket list. Every line of the
scorecard's output matches, except three. The week of 31 August's tokens per change are 11.0
million, not 11.4, because transcripts have aged out. The unfinished last week gained 3.7
working hours. At today's origin/main September's rank correlation is 0.21, not 0.23. Both
figures regenerate byte-identical.
| Paper's figure | Re-run | Verdict |
|---|---|---|
| 47 a week since mid-July, 95 "in June"; 131 PRs a week | 46.5, 94.8, 131.3 | reproduce; the June weeks are 8 to 29 June, below |
| 998 mature; 14 / 31 / 78 / 114 / 139 not; 763 pass, 76.5% | identical | reproduce; the definitions, below |
| "Three in four merged tickets" | 76.5% of the mature population; 71.6% since mid-July | true only over the whole population |
| 66 escapes name the change; median filed the same day, 86% in 7 days | identical | reproduce; 35 of the 66 name the finder, below |
| Blocks: 152 / 232 / 127; 35.4 / 25.7 / 48.8 dispatches and 4.43 / 2.67 / 2.79 hours a change; medians 13 / 9 / 18 | identical | reproduce; July's hours, below |
| Tokens 11.4 / 14.4 / 11.4 / 16.0 M a change | 11.0 / 14.4 / 11.4 / 16.0 | transcripts age out |
| By repo, BLOCKED entries, correlations, process weight, docs-only 43 | identical | reproduce |
| CVs and detectable ratios, lag-1 −0.13 to 0.19 | identical | weekly rows hold; per-change rows do not, below |
| "A cap of 1 to 4 hours gives 3.2 to 3.8 hours a change" | 3.0 to 3.5 pooled, as the table pools | the paper averaged weekly ratios instead |
The definition of correct inherits the finder rows, and more of them than survey-check.md
found. The scorecard reads introducedBy from the reliability verdicts, which survey-check
found often names the ticket whose review found an older fault. By their own reasons
("pre-existing", "predates", "older code", "left out of scope"), 35 of the 66 escapes that name
a change are finder rows. They include all 18 that survey-check listed. The scorecard's window
opens two days before the merge, so it also re-admits 11 negative-lag rows that the
reliability script drops, 10 of them finder rows. The other 31 were filed a median of 3.3
days after the change's last merge, 71% within 7 days and 97% within 30. So "the median is
filed the same day … this agrees with the reliability paper's 1.4 days" is wrong twice. The
real lag agrees with survey-check's corrected 4.1. Eight of the 31 changes counted incorrect
are named only by finder rows (LIN-1815, LIN-2037, LIN-2291, LIN-2331, LIN-2333, LIN-2351,
LIN-2354, LIN-2384). Without them 770 pass, not 763.
The named-fix test matches mentions. A later ticket counts as a fix of a change when its
subject or its ticket's title says fix, and the change's id or PR number appears anywhere in
its description (survey-reliability-git.mjs:86-98). 36 of the 78 say fix only through the
title, 69 name the change only in the description, and a PR number is not checked against its
repo. One reader read the text around each mention, not blind. About 11 clearly blame the
change, 15 are unclear, and about 52 are mentions: "prior history … is unrelated", "do NOT fix
this", line-number shifts, nine tickets named in one plan. With those and the finder rows
corrected the mature pass rate would be about 82%, not 76.5%. This reading is one reader's and
first, so version 2 reports it as a limit, not a figure.
Completeness is biased both ways, not only down.
- Omitted origins. It takes a follow-up's origin to be its parent, even when the parent is
an epic. 38 mature changes named in the text of such follow-ups count as complete. It ignores
kind:review-residuefilings. Filings that say "Filed by LIN-X close-out" or "Routed from LIN-X review" name 78 more mature changes (sampled, not all read). - The 23-of-40 rate does not transfer. The Limits cap the downward bias with
never-worked-pile.md's 23 of 40, a sample of never-worked follow-ups. A quarter of follow-ups were worked. - Right-censoring. Completeness has no window at all, so recent weeks are high on both counts. The week of 21 September shows no named fix and 6% incomplete, against 8% and 20% in the mature weeks since mid-July. At mature rates it would have about 50 correct, complete changes, not 65, "not a few".
The 95 → 47 fall: June's figure is the weeks of 8 to 29 June, and most of the fall is fewer
merged tickets. survey-scorecard.mjs:274 averages the weeks beginning 8, 15, 22 and 29 June;
the last runs into July. June's four calendar weeks, 1 to 22 June, give 69 a week. In the week of
1 June only half of merged PRs named a ticket, against 81–100% in every week after, so 69 is low
too. On the paper's weeks, merged Done tickets fell from 111 to 64 a week before any filter. Of
the 48-change fall, 37.5 (78%) is fewer merged tickets and 10.8 (22%) a lower pass rate, 84% to
72%. The pass rate fell for reasons June could not show:
- Finder-row escapes: 4.3 points; none are in June.
- Follow-ups: the
kind:follow-uplabel came into use around 10 June. - Routed Bugs: June has none.
- What did not change: the join does not lose later tickets, and the median change is about 80 production lines in both periods.
why-throughput-halved.md (LIN-3155), which landed while this check ran and is not checked
here, finds the same four-fifths and one-fifth on the same weeks.
The weekly detection power holds; the per-change power does not. The method is a z-test on log means, exp(2.8·s·√(2/k)), with weeks independent and s from 11 weeks. Taking t-quantiles for an s estimated on 10 degrees of freedom, the weekly count needs ×2.66 at four weeks and ×2.0 at eight: "about eight weeks" is the edge. The per-change rows assume about 47 independent changes a week, but only 41 a week have runner data. Their weekly means vary 2.7× (dispatches) and 2.1× (hours) more than independence predicts (F = 15.0 and 7.3 on 10 and 444 degrees of freedom). On the observed spread of the weekly mean, four weeks detect ×2.8 in dispatches and ×2.0 in hours, and eight weeks ×2.1 and ×1.6, against the paper's ×1.4 and ×1.3. A halving is not visible in one to two weeks (×7.9 and ×4.3). The per-change series is also a different measure from the headline one. Its dispatches are those mapped to each ticket, not fleet dispatches over changes, and its hours are the runner's uncapped clock. The paper's conclusion survives only weakly: hours per change sees a doubling at about four weeks, the weekly count at about eight.
The scorecard's weekly dispatches are clear of the July run-log undercount; three July figures are not. The fleet series starts on 13 July, after the 6–11 July gap, and 17,805 of the 17,910 oplog items since 12 July are in the run logs. What the undercount and the scope touch:
- Hours in the July block. The 673 working hours in the July block are every oplog session. The dispatches are Harbour's workspaces only. About 19 of those hours are from other workspaces, and 94 are from sessions the run logs cannot place. On Harbour's own sessions July is about 3.7 hours a change, not 4.43, so the fall "from 4.4" is overstated.
- July's rank correlation, 0.06, includes tickets merged 1 to 12 July, before the oplog, whose median is 5 dispatches. From 13 July it is 0.11.
- The per-ticket medians and the per-change noise use
survey-effort-runner.mjs's every-item-ever mapping. From 13 July it barely touches pre-gap logs (0 of 178 tickets), and the July median of 13 is not low.
Two citations do not stand after survey-check.md, and version 2 replaces both. "Per
ticket, dispatches doubled while working time held" is about 1.4× from July, almost all warm
beats. The paper's own blocks agree: the median went 13 → 18 and fleet dispatches per change
35.4 → 48.8, both 1.38×. "Nearly doubled from August" is true only because August is the cheapest
block. And the risk gap does not "disappear once size is held fixed". The data cannot tell equal
effort from a difference of half either way.
Smaller slips, each corrected in version 2. The five changes with no state did not merge after the census. They merged 13 to 28 June: LIN-450 and LIN-704 are Done but missing from the team list, and LIN-756 to LIN-758 no longer exist in the tracker. "66 of 190 escaped Bugs name the change" is about 31. The other citations land where the paper says.
The lines of docs/steady-base.md that change
The anchor (at fd6b1352) was written from version 1 of all three papers. This check did not edit
it. Nineteen lines change because of this check, and one more because it misquotes
survey-check.md:
| Line | Now | Should read |
|---|---|---|
| 23 | "Without that layer, supervision was flat at 26–28%." | 28%, 24%, 15% and 26% by week: flat to falling (survey-check.md) |
| 29 | "A third of wakes change nothing, and they are 30% of the supervision bill." | 31% of wakes, and 25% of the bill; the Runner's 94% stands |
| 30 | "Most of the answers are already held in deterministic code …" | Much of it is (the completion gate, wake delivery, stall handling); how much is unmeasured. The prompts do not ask for polling alongside wakes, and part of the Done gate is in code |
| 31 | "24 of the 26 on record were in mechanical actions, and 21 of those were bugs in the plumbing." | 22 of the 25 on record, on a reading of each ticket; 21 were in code |
| 41 | "… make nearly all the catches that matter." | Behavioural tests alone kill most sampled logic mutants; in sessions the 225 production edits are a ceiling, and a reading of ten found one catch |
| 42 | "turned CI red only twice since June … a pin demanded a bump about as often as it preceded a fix, and 18 tickets since June exist mainly to bump one." | Twice as unit tests; count pins inside browser specs three more times. Bumps at least as often as catches, on one or two catches a class. 18 tickets (20 by the paper's rule) exist mainly to repair or bump a pin; about five are bumps |
| 46 | "four files idling about 25 seconds each on an open keepalive handle." | … on a 25-second timeout timer in routes/proxy.js that is never cleared; it also costs about 20–30 s of parallel wall-clock |
| 75 | "(95 a week in June; …)" | 95 a week in the weeks of 8 to 29 June; 69 in June's calendar weeks, when half of PRs named no ticket |
| 76 | "26–49 dispatches and 2.7–4.4 working hours per correct change" | July's 4.4 counts other workspaces' sessions; on Harbour's own it is about 3.7 |
| 79 | "The instrument detects a 1.4× shift in dispatches per change, or 1.3× in hours, within four weeks." | ×2.8 in dispatches and ×2.0 in hours over four weeks; hours per change sees a doubling at about four weeks, the weekly count at about eight |
| 96 | Row 1: "A third of wakes change nothing, 30% of the supervision bill"; "stop prompts asking for polling on top of push wakes"; "(12 of 26)"; "~10%" | 31% of wakes, 25% of the bill; the prompts do not ask for that polling; 15 of 25; the ~10% was 30% of the 35% share, and 25% of it is about 9% |
| 97 | Row 2: "re-arms on 94% of its wakes" | 94% of its wakes change nothing |
| 98 | Row 3: "the answers are already in code" | much of it, unmeasured; 77% mechanical stands, 86% on a fresh blind recode |
| 99 | Row 4: "(the open keepalive handle)"; "up to about a third of serial unit time" | the uncleared withTimeout timer; a third of serial time and about 20–30 s of parallel |
| 101 | Row 6: "18 bump-only tickets since June, 10 in September; no CI catch on record"; "~2–3 tickets a week of bump work" | about five of the 18 are bumps; e2e count pins failed PR CI three times; the weekly rate rests on the 18 |
| 121 | "fall from 95 a week in June to about 47 from mid-July" | from 95 in the weeks of 8 to 29 June; about four-fifths of the fall is fewer merged tickets, a fifth a lower pass rate |
| 123 | "The research papers themselves are unreviewed by a second document for this wave" | checked by survey-check-2.md; each has a version 2 |
| 136–138 | The three papers, unversioned | each "(v2)", and a row for survey-check-2 |
Unchanged and confirmed: line 27's 77% and 27%, line 45's 77% outside any test, line 44's flakes, line 47's stale mutants, and line 77's tokens.
Method
Each paper's own commands were re-run unchanged. Outputs went to the git-ignored data/, and
figures to scratch directories; no committed figure was redrawn.
# what supervisors do: a fresh extract over the paper's window, then every step; and over a copy of the author's cache
node scripts/survey-supervise-extract.mjs --since 2026-09-01 --until 2026-09-30T12:00:00Z --out data/check2/sup-fresh
node scripts/survey-supervise-classify.mjs --in <dir>/steps.jsonl --out <dir>/coded.jsonl
node scripts/survey-supervise-sample.mjs --dir <dir>
node scripts/survey-supervise-failures.mjs --dir <dir> --offline
node scripts/survey-supervise-analyse.mjs --dir <dir>
node scripts/survey-supervise-render.mjs --dir <dir> --out <scratch>/fig
# the recode's sample (committed with this check)
node scripts/survey-supervise-pps.mjs --dir <author's cache> --n 160 --start 0.37
# test estate: shape and validation at HEAD, timing at c65b7dd8, CI and sessions fresh, friction at c65b7dd8
node scripts/survey-tests-shape.mjs --rev HEAD --sample 70 --seed 3151 --validate
node scripts/survey-tests-timing.mjs lv # in a scratch worktree at c65b7dd8
node scripts/survey-tests-ci-runs.mjs; node scripts/survey-tests-ci-fetch.mjs; node scripts/survey-tests-ci-classify.mjs
node scripts/survey-tests-local.mjs; node scripts/survey-tests-friction-git.mjs; node scripts/survey-tests-friction-tickets.mjs
node scripts/survey-tests-analyse.mjs; node scripts/survey-tests-figures.mjs
# measuring throughput: git, runner, fleet and tokens fresh; the scorecard over the author's snapshots
node scripts/survey-effort-git.mjs; node scripts/survey-effort-runner.mjs
node scripts/survey-growth-fleet.mjs --json > data/survey/scorecard-fleet.json
node scripts/survey-effort-fleet.mjs --since 2026-08-31 --until 2026-09-30T06:50:00Z
node scripts/survey-scorecard.mjs --cap-hours 1|2|3|4 --svg <scratch>
Caches. The three authors' caches survived in other session workspaces on this machine and were copied, not trusted. The supervisor extract was compared with a fresh one (identical), the test estate's CI and session snapshots with fresh pulls, and the scorecard's ticket list with a fresh list read.
The recode. scripts/survey-supervise-pps.mjs orders the census by session start and step
and walks it in steps of one-160th of the total weighted tokens, from a fixed start of 0.37 of a
step. It skips the paper's 220 steps; a hit moves to the session's next step, which happened
twice. Cards are built exactly as the paper's sampler builds them. The two coders were given
only the paper's codebook and the cards. Coder 2 read them in reverse. Neither saw the paper, its
codes, the rule classes or the other's file. Agreement is raw proportion and Cohen's κ. A step
scores 1 if both coders marked it M, 0.5 if they split, 0 if neither. The mean score over the
sample is the token-weighted share, with a 20,000-resample bootstrap interval. The codings, and
each step's rule class added afterwards, are in survey-check-2-codes.json.
Code and tickets. Every code line a paper cites was read with git show <sha>:<path>, and
every uncited claim about code was checked with git grep in both repos. The idle was traced
with a preload that wrapped setTimeout, and tested by unref'ing each candidate timer in a
scratch worktree. About 90 proxy reads, no writes except this ticket's status and comment. Each
was paced at one per 8 seconds, because a sibling session shared the proxy's 60 a minute:
- all 26 supervisor failures, a systematic 15 of the 57 excluded candidates and 8 targeted;
- the 18 pin tickets and 2 misses;
- the ticket list and 6 details for the scorecard.
Hand readings. Each was one reader, not blind:
- all 33 census and text-pin session episodes;
- every 16th behavioural production-fix episode;
- every 9th red PR attempt;
- the text around each of the 78 named-fix mentions;
- the reasons of the 66 escapes.
The one-off analyses that produced the sensitivity, cycle, power and decomposition figures are in the session scratchpad and are not committed; every number they give is stated here with its population.
Limits
- Every reader shares a tier with the authors. The re-runs, the recode and the readings are independent of the authors' sessions and reasoning, not of their model tier. The recode uses the paper's own codebook, so it tests the coding, not the definition. The definition is where the mechanical share is exposed.
- The recode's agreement on class is high enough to be suspicious. 98% on class, κ 0.98, is higher than the paper's 94%. Both coders read quickly, and a card's wake text makes many steps easy. M or J agreement, κ 0.68, is the figure to weigh, and it is lower than the paper's 0.83.
- The PPS sample's precision. 160 steps give ±5 points on the share. The sample's rule mix matches the census, but its judgement classes are thin: 17 judge steps and 11 beats.
- The failure re-reading, the named-fix reading and the pin-episode reading are one reader each, not blind. They show where the papers' labels are weak. Counts from them are first readings, and version 2 reports the named-fix reading only as a limit.
- The parallel-time figures are a simulation. Measured full-suite runs on the runner machine, under heavy sibling load, pointed the same way but varied by 20 seconds.
- The effect on the fleet. The check re-ran a simple-dispatcher script that, run as a module, starts the mutation gate. It was stopped mid-mutant, and the one mutated file in this session's checkout was restored and verified clean at 3b1e734. The live runner's checkout was not touched.
- Transcripts age out daily. The supervisor extract was still identical today; a later re-run of this check will not be.
Next
- Which of the scorecard's named fixes are fixes, and which follow-ups are the change's own?
The correct and complete tests rest on text matches: 69 of 78 named fixes are a mention
anywhere in a later description, and 116 mature changes are named as origin in filings the
scorecard ignores. A blind second reading of both sets would turn the scorecard's correct and
complete rates from bounds into estimates. This goes into
proposals.md. - The existing
proposals.mdline on per-change sensitivity quoted ×1.4 at four weeks; it is corrected to the observed spread, ×2.8 in dispatches and ×2.0 in hours. why-throughput-halved.md(LIN-3155) landed while this check ran. It uses the same "June" weeks and, for its escapes by tier (9.6% against 2.4%), the sameintroducedByfield, so its check should test how many of those escapes are finder rows.- Each paper's own Next stands. The test estate's "what holds them?" is answered here. The supervisors paper's first Next, replaying gate replies against the runner's own state, is the measurement that would say how much of the mechanical share state alone decides.