Do the prototype-concepts and what-hides-between-sessions papers hold up?
Their scripts do, and their direction does; several of their figures and readings do not. Every committed script in both papers re-runs from the authors' snapshots, and every table reproduces: what-hides-between-sessions' byte for byte, prototype-concepts' with five medians moving by under half a percent. Both headline answers stand. What fails is in what each figure is taken to mean.
prototype-concepts. Its prototype figures are read correctly at the pinned commits, but four readings go further than their sources, and its cost comparison is not like for like in a way the paper misreads:
- The per-piece costs differ by price row, not by the shape of the work. The fleet moved from the older frontier price row to the newer, cheaper one in the week of 22 September. The one-session papers and checks ran 90–100% on the newer row, the pipeline papers 77% on the older one, and the code changes half their list price on it. The paper's "a multi-session paper costs 1.7 times as much per weighted token", and the Next item and proposal built on it, are that switch. Within one row a paper, a check and a code change cost the same per weighted token to within 3%.
- As ratios, a Lighthouse study is 1.35× a one-session Harbour paper at list prices and about 1.0× in weighted tokens; 0.76× and about 0.57× a paper with its share of the check. A code change is 2.61× a paper at list prices, 1.70× in weighted tokens, and 1.97× priced wholly on one row. The Lighthouse rise from early to late studies is 3.7×, but the reader review itself is a small part of it, so option E's "3.7× after adding it" does not hold.
- The tag-two 51% against 29% is a contrast its own repository calls "not established". CodeWiki "did not improve" rather than "got worse"; tangle's token premium is a median 1.44×, not 2–5×; harbour-evals' review and research failures are put down to fragile fixtures by its own report.
- Lighthouse. Its secondary frame adds 21 post-release changes, not 29. "Not from its second agent" is contradicted: review agents made 15 of the 35 post-release corrections, in rounds the keeper opened. Its bound on subagent spend was exceeded four times, not three.
- The checks coding counts every claim a check found wanting as "load-bearing". By the checks' own word, about 20 of 24 documents had a load-bearing claim corrected, not 23. One check ran before merge.
what-hides-between-sessions. Its register, its scoring and its nine hits all hold: each hit is the incident, and each lead time recomputes within 0.1 hours. Six claims do not hold:
- "About 90 real alarms nobody recorded" counts alarms real by class. A systematic hand check of 33 found 20 real, 7 harmless blips and 6 false, which puts the distinct real faults at about 40 (25–50). D1's unrecorded bursts were nearly all errors that healed in minutes.
- "About half" of the incidents on record is 9 of the 19 the windows scored. Counting the five credential incidents in D1f's own window it is 9 of 24. Option A's two rules caught 7 of 19, 0–9.3 hours early, not "9 of 19, 0–21 h".
- The 108 hours of lost-wake idle. 66 of them are waiters the runner had already completed, mostly after their ticket was Done. 42 hours are sessions still waiting.
- The failsafe re-fires' 1.6% "outside P4" double-counts map row 1 by 0.9 points. The quiet re-fires into supervisors are in held-or-fresh's class route.
- Option B's "23 lost wakes" are 23 failed completion posts. At least 4 of their parents were woken anyway, and four failed because their items no longer existed.
- The 23 September autopilot that "sat 23.7 hours" sat because the dispatcher was down for a day. That outage is on no ticket.
The median of 16 hours to discovery reproduces, but it is driven by onsets known only to the day. Where the record gives the onset's time, the median is 3 hours.
Reconciled against the three earlier papers, the lost-wake record agrees: 11 of 15 in the
runner. The detectors do not double-count rows 1–3 in tokens once the re-fires are moved. They do
overlap row 3 in mechanism, since row 3 moves the liveness clocks into code. Both papers are at
version 2 in this PR, with this check's author added. Ten lines of docs/steady-base.md would
change and three rows be added; they are listed below. This check does not edit the anchor.
Findings
Every script re-runs, and every printed number reproduces.
- what-hides-between-sessions. Re-run from the author's snapshot, the analysis, detector and wait-coding outputs and both figures are byte-identical. The runner detector, which reads the live oplog, moves only two bookkeeping counts by one or two lines. A fresh extraction from the transcripts, cut to sessions whose first turn came before 10:14Z, gives the paper's 2,076 sessions and every figure but two. September's tokens rise 0.1%, because four long sessions are still running into October. The failsafe count rises by one, because the runner script dates the newest run log's fires at its start.
- prototype-concepts. The Lighthouse, Harbour and figure scripts give the author's JSON and
byte-identical charts. The token and dollar censuses re-run with
--untilgive the same 2,004 sessions and totals. The classes script reads origin/main as it stands, so at today's head it adds this wave's own three papers to the one-session population: 33 at 0.97 of the median. Pinned to d88e2216 it reproduces.
| Figure | Version 1 | Checked | Verdict |
|---|---|---|---|
| prototype-concepts | |||
| Pith: 236 against 286 rubric points, 0 wins in 15; "lost on 14 of 15" | 0 wins, 14 losses, 1 tie; the judge takes the exploring agent's answer as ground truth; about 14 runs, nearly all lost | holds | |
| CodeWiki: 6 of 15 accurate to 4 as pages grew; "got worse" | 6, 3, 4 of 15 in one run, which the project blames partly on parsing bugs | corrected to "did not improve" | |
| Browser-agent 3/20 → 19/20 → 20/20; 29/40 → 39/40; tuned on a 0.6B model | reproduce; 20 trials an arm, after-arm repeated; the title tool carried over to a 1.7B model | holds | |
| "No prototype produced better than a single-run controlled comparison" | not literally true of Browser-agent or of the pointer probe's three runs | corrected | |
| tag-two: 24 of 47 with the record against 4 of 14 without (51% against 29%) | 11 graphs against 3, another commit, a model check that varies by a node a graph; the source calls the contrast "not established" | corrected: about half, with the record | |
| tangle: 14 of 18 found more, "often at 2–5 times the tokens" | 14 of 18 hold; on those 14 a median 1.44×, 5 at 2.7–5.4×, 2 at fewer | corrected | |
| harbour-evals: review and research "failed for every model" | they did; its report calls them fixture fragility, and the review fixture has no filename to cite | corrected | |
| Strength codes | Lighthouse's staged review (2) and harbour-evals on research cost (2) are generous at 1; the Limits' proposal to lower Pith and tangle misreads the scale | Limits corrected | |
| 12 relevant + 6 not + 2 empty = 20 repositories; 14 concepts; the five pins unchanged | reproduce | holds | |
| Lighthouse study medians, early (7) and late (6), and over all 13 | reproduce exactly | holds | |
| "LH011–LH017 (reader review added)" | four of the six (LH012, LH014, LH015, LH017) | corrected | |
| One-session paper n30; check n11; pipeline n8; code change n202 | reproduce | holds | |
| Check "covers 2–3 papers" | three cover one each | corrected | |
| Paper with its check share, n22 | 21 papers and an essay; two checks have no ticket. Median 0.1% lower, IQR ends ±0.4% | corrected | |
| LinearViewer and simple-dispatcher code medians; size bands | four medians move 0.03–0.3% | corrected | |
| A check is a third of a checked paper (33%); option D's "about a third" | reproduce (2.95M weighted a paper) | holds | |
| The step from paper to checked paper | "its share of the check" | the checked papers' own median is 1.20× all 30 papers', before any check | qualified |
| "All 16 dispatches with token records re-price to the cent" | true, but the same formula re-applied; 3 of 16 belong to a costed study | qualified | |
| "Studies after 26 September carry only transcribed figures" | only LH002 of the 13 has token records | corrected | |
| Lighthouse's bound on subagent spend a session exceeded three times | subagent spend at 1.5× the bound (the improvement round, whose quoted figure included its driving session), 1.8× (LH015, against a raised bound) and 1.9× (LH016, omitted), plus a driving session at 2.2× that the bound does not count | corrected: four | |
| Option E: "3.7× after adding" the reader review | medians rose 1.3× in research, 3.9× in review, 8.7× in driving sessions (partly a change in how they were measured); subagent spend 2.1×; the reader pass itself 0.2–0.8× a Harbour check | corrected | |
| "Top-priced tier at 2.5 times the frontier tier's list rates" | 2.5× on input and output, 1.25× on cache reads; three review dispatches cost 2.2–2.3× the same tokens at the frontier row | corrected | |
| One-session against pipeline paper: 2.3× at list prices, 1.3× weighted; list price per million weighted tokens 1.7× as high for the pipeline paper | reproduce (1.73× by medians); the cause is the price row, not cache structure: 1.50× on the newer row, 1.23× on the older | corrected; Next and proposal replaced | |
| LH016: 66 changes, 42 before and 24 after release; evidence review 23; reader review 17; recheck 10; 15 of 17 in passed passages; replays 5, 7, 4 of 16 | reproduce | holds | |
| LH016: "plus 29 more after release in its secondary frame" | 21: amendment 1 counts 8 of the 29 in the primary frame only (tally.txt:77, 45 = 24 + 21) |
corrected | |
| "Not from its second agent" | review agents made 15 of the 35 post-release corrections, in keeper-opened rounds (LH016.md:69) |
corrected | |
| "2 carry correction notices and 5 later-evidence notices (13 notices in all)" | 3 notices on 2 articles and 10 on 5 | corrected | |
| 23 of 24 checked documents had a load-bearing claim corrected; 95 in all; about 9 figures at the median | 23 had a claim corrected; by the checks' own word "load-bearing", about 20 (growth-atlas, why-legs-repeat and learning-while-the-tools-change drop out); 95 counts every failed claim; the median is 4 claims and 9 figures. A blind recode of 8 pairs agreed on the answer in 7 | corrected | |
| "Harbour's checks run after a paper merges" | 23 of 24; the check of what-should-an-agent-leave-behind read the draft and found nothing | corrected | |
| Headline held in 20, changed in 4; 23 of 24 at version 2 or later; 44 faults in 18 of 100; escapes 5.7 and 27.3 per 100 | reproduce | holds | |
| what-hides-between-sessions | |||
| 107 incidents "of nine patterns"; 92 invisible; finders 28/26/19/13/9/3; repos 54/44/9; every pattern row | reproduce; 6 of the 107 are in none of the nine | corrected | |
| Median 15.8 h to discovery (headline "16"; option A "10–15"; Next "about 15") | reproduces; 3.0 h where the onset is timed, 19.0 h where it is a date; 47 of 64 onsets are dates, not 65 | qualified | |
| 9 hits; leads 0.8, 8.8, 9.0, 9.3, 4.9, 5.1, 21, 0, 5 min; median 5.1 h | each alarm is the incident; leads recompute within 0.1 h | holds | |
| "About half" caught: 9 of 19 | 9 of 19 scored; 9 of 24 with D1f's own window; D5x's window holds seven July–August lost wakes never scored | corrected | |
| Option A (D1 and D2): "9 of 19 incidents 0–21 h sooner" | D1 and D2 caught 7 of 19, 0–9.3 h; the 21 h is D5x's | corrected | |
D1: "20" unrecorded; "19 more LINEAR_AUTH"; "21 bursts" in 2–25 September |
20 = 17 auth + 3 transient; 21 auth bursts in the month, 14 in 2–25 September. LIN-2993 is scored as a hit in the file and a miss in the text; as a miss D1 has 23 | corrected | |
| "About 90 real alarms" (20 + 43 + 3 + 22 + 5 = 93) | real by class | a sample of 33: 20 real, 7 blips, 6 false; about 40 distinct real faults (25–50) | corrected |
| False alarms: D1 16, D2 2, D5 0, D5x 0, D7 2 | under the scripts' rules 16/42, 2/57, 0/4, 0/23, 2/7; D5 and D5x set to 0 in code. In the sample: D1 1 of 7, D2 3 of 11, D5x 2 of 7, D7 0 of 5 | qualified | |
| "Of September's 27 lost wakes, 5 were rescued by nothing" | 2; three were woken within a minute | corrected | |
| D5: the 23 September autopilot "never woken", 23.7 h | it sat; the dispatcher was down 06:53Z 23 September to 06:35Z the 24th, on no ticket and not in the register | corrected | |
| D7: duplicates on LIN-3125, LIN-2667, LIN-2718 | all real; two more on LIN-2560 | holds | |
D5x: 22 of 23 still wrote hook.done_posted; no retry |
reproduces at hook.js@3366748:2058-2065 and feedback.js:41-70 |
holds | |
| Option B: "23 lost wakes since July, about 8 a month" | 23 failed terminal posts, 8 in September; at least 4 parents woken anyway; four on 26 July failed because the items were gone | corrected | |
| P2 in September: "1 cycle and 2 recorded orphaned waits" | the script counts 3 recorded (LIN-2559, LIN-2630, LIN-2932) and the cycle is 1 October's | corrected | |
| P5: 14 woken "24–50 min late" | 13 at 24–47 min, one 5.3 h | corrected | |
| P5: 108 h of lost-wake idle | reproduces; median 5.4 h, 10 at the 6 h cap; 66 h in waiters the runner had completed | qualified | |
| P6a: 219 fires on 149 sessions; 165 h; 130 of 142 re-fires worked | 206 stall-failsafe and 13 watchdog fires, one misdated from 1 October; 165 h is logged silence over the 80 fires from execution; "worked" means at any later time | qualified | |
| P6a: 54.8M of re-fires (1.57%) "outside P4" | 30.4M of it is quiet re-fires into supervisors, inside P4 and map row 1; 0.70% outside | corrected | |
| The periodicals: 42 reports, 13 of 15 templates, 67 runs, 109 follow-ups, one first find, 9 of 12; 112M (3.2%) | reproduce; the 26 September batch alone is 109M (3.1%) | qualified | |
| "About 40 seconds of CPU" | 35 s without the wake extract the runner detector reads, 55 s with it | corrected | |
| 2,076 sessions; transcripts to 10:14Z; oplog to 10:06Z | the snapshot runs to 10:36Z (2,077) and 10:33Z; 2,076 is the sessions begun before 10:14Z | corrected |
The cost comparison is like for like on the frontier row only for Lighthouse against Harbour's
papers and checks. Lighthouse prices every dispatch from its own prices.json. Prototype-concepts
priced Harbour's transcripts from the same table, each message at its own model's row. Three things
differ between the pieces:
- The rows. The table has two frontier rows. At list prices a weighted token on the newer costs
0.53–0.55 of one on the older, because the newer is cheaper on every token and charges a cache
read at a twentieth of input where the older charges a tenth. The weighted unit charges both at
- September's fleet ran 65–72% on the older row each week until 22 September and 77–98% on the newer one after. Lighthouse ran wholly on the newer row, plus its top tier for review.
- The tiers. Lighthouse reviews on its top tier. That costs 2.2–2.3× the same tokens at the newer frontier row on its three review dispatches with records, and about 2.7× per weighted token, because the weighted unit counts the top tier at 1. Harbour's checks ran wholly on the newer frontier row.
- What a piece includes. A Lighthouse study includes its driving session (self-reported, and measured at commit time before 28 September), research, writing, review, its post-release rounds and LH016's replays. It leaves out the keeper's relayed review, the improvement round and a website build. A one-session Harbour paper is its one session and its subagents, with no supervisor. A code change is every session entered on its ticket, about a quarter of its cost in the ticket's own autopilot and stepper, and none of its epic's supervision.
Read on one row, the ratios are: a Lighthouse study about 1.0× a one-session paper in weighted tokens, with its premium in the top-tier review; a code change 1.97× a paper. Lighthouse estimates output tokens as characters over four. On Harbour's own transcripts that gives about a third of the recorded output, which would put Lighthouse's studies about a quarter low. This was not measured on Lighthouse's transcripts.
The prototypes are read fairly against their own repositories, with four overreaches. Every clone's head is the sha the matrix pins, and the five pins the earlier register used are unchanged. The overreaches are in the table: tag-two's contrast, CodeWiki's direction, tangle's premium and harbour-evals' failures. Two codes on the paper's own 0–3 scale are generous: Lighthouse's staged review is a tally of who changed what, with no comparison process, and harbour-evals measures nothing about research cost. Browser-agent's 2 is if anything harsh.
"Load-bearing" in the checks coding means "failed". prototype-concepts-checks.json counts
every claim a check found did not hold. Where the checks draw their own line, the count drops:
survey-check.mdnames four load-bearing failures across three papers, none of them growth-atlas's, but the coding gives that check 13;survey-check-6.mdcalls why-legs-repeat's five failures secondary;- the check of learning-while-the-tools-change ran in the session that filed the essay and found failures "of reach and of position, not of fact".
A blind recode of every third pair (8) agreed on the headline answer in 7. It agreed on the load-bearing count within one in 5.
The detectors' hits are real; most of their unrecorded alarms are not distinct real faults.
Before reading any alarm, "real" was fixed as a fault that idled a waiting session 15 minutes or
more, lost or duplicated work, failed an action a retry did not recover, or needed a person or the
failsafe to recover. A harmless blip is a fault that healed itself in about 15 minutes, with
nothing parked or lost. The sample was every k-th unrecorded alarm by time, and all of D5's and
D7's (survey-check-10-alarm-codes.json):
| Detector | Unrecorded | Sampled | Real | Blip | False | Estimated real |
|---|---|---|---|---|---|---|
| D1 shared-error burst | 20 | 7 | 0 | 6 | 1 | about 1 (0–7) |
| D2 stopped or circular wait | 43 | 11 | 8 (2 usage limits) | 0 | 3 | about 31 (19–39) |
| D5 lost wake | 3 | 3 | 3 | 0 | 0 | 3, none distinct |
| D5x lost terminal post | 22 | 7 | 4 | 1 | 2 | about 14 (7–17) |
| D7 duplicate launch | 5 | 5 | 5 | 0 | 0 | 5 |
That is about 54 before duplicates are removed. Duplicates are the same event caught by two rules,
such as a lost [done] that D2 and D5x both see. Without them, about 40 distinct real faults
remain (25–50), of which about 8 are usage-limit stops. D2's false alarms share one mechanism: when
a wake lands within three minutes of the waiter declaring its wait, the rule's three-minute burst
merge swallows the wake and counts a lost wake up to the cap. D5x's join to a parent missed real
parents three times, so its "no parent" rows are not harmless either.
The register's counts all recount exactly. A second reading of 15 rows, taken systematically, agreed on filing time and repo in all 15. It would recode two finders (LIN-800 and LIN-2232 were found by review) and two onsets (LIN-1113, LIN-1815), and argue five pattern and three invisibility calls. Three incidents are missing from the register:
- the dispatcher's day-long stop on 23–24 September;
- the 22–23 September dead-credential night, on record in LIN-2933's comments and LIN-3018;
- the periodicals' four-week outage, which the paper treats as their only first find.
These change no headline. The version 2 paper states them in its Limits rather than recoding the author's register.
Reconciled against the three earlier papers, the records agree; the token savings overlapped.
wake-inventory.mdv2 puts 15 of the 25 supervisor failures on record as lost wakes, 11 on the runner's side. The register codes the same 15 as runner 7, both 4 and Harbour 4, and two of them (LIN-881, LIN-2932) as circular or orphaned waits, not lost wakes. "Runner or both" is 11, so the paper's "11 of 15 in the runner" agrees. The register's 20 lost wakes are 14 of the 25 plus six others. One of those six, LIN-2468, is a casewhat-supervisors-do.mdv2 set aside as latent, with no incident.what-supervisors-do.mdv2's 25 failures are all in the register. Two of the nine detector hits are among them (LIN-2511, LIN-2932).held-or-fresh.mdv2's 11 failures that would not exist under a relay are 11 of the register's 107: ten lost wakes and LIN-881. One detector hit (LIN-2511) and one miss (LIN-2720, an opencode session with no transcript) are among them. A relay and the detectors therefore claim the same incident once in nine hits.- Do the detectors' savings double-count map rows 1–3? In tokens, partly. The paper puts
failsafe re-fires at 54.8M, 1.57%, "outside P4". But 270 re-fire deliveries make that figure. The
quiet ones into supervisors, 96 deliveries and 30.4M, are inside held-or-fresh's 3,233 quiet
wakes, and its class route sends failsafe re-confirms and silence re-fires to code
(
survey-check-7-held.mjs:68). That is map row 1's −5.7%. Outside P4 the re-fires are 0.70%, and the tokens outside quiet wakes are about 1.6%, not under 3%. Flood tokens are inside P4, as the paper says. In hours there is no double count, because the anchor sizes rows 1–3 in tokens only. In mechanism, option A and option B overlap row 3, which moves "the completion gate, re-arming, restating and liveness clocks" into code (steady-base.md:172). The paper's "it does not overlap map rows 1–3" holds only for tokens.
The options' sizes, after correction, and where they overlap. No option is added here.
| Paper and option | Size as corrected | Overlaps |
|---|---|---|
| prototype-concepts A, a short pointer | ~1.5–2% (as row 14) | it is anchor row 14 and starting-context B |
| B, code holds the sequence | up to ~27% with rows 1–2 | it is anchor row 3; the same lever as what-hides A and B from the other end |
| C, a done-and-answered list | inside −9% to −19%; tag-two supports "about half repeated with the record", not the contrast | inside anchor row 16 |
| D, check before merge | about cost-neutral, 2.95M weighted a paper; the one pre-merge check found nothing | no anchor row |
| E, a reader pass | the pass itself 0.2–0.8× a Harbour check, not 3.7× a study | no anchor row; costs what D saves in waiting |
| F, a repeated fixture eval | enables where-judgement-happens option 3 (2–5% of ticket cost); the 7–14 spread was across harness revisions | no anchor row |
| G, graph investigation (not recommended) | more coverage at a median 1.4× the tokens | none |
| what-hides A, live D1 and D2 | 7 of 19 on record, 0–9.3 h sooner; up to 42 h a month of supervisors waiting; under 1% of tokens; D1 needs a harm or duration threshold | row 3's liveness clocks in mechanism; prototype B |
| B, a truthful completion post | 23 failed posts since July, 8 in September; at least 4 woken anyway | row 1's wake-delivery safety check; row 3 |
| C, periodicals read the alarm log | up to 3.1% a batch | none |
| D, retire vigilance text | a few KB of 36 KB | anchor row 11 (freeze prompt sizes; lessons land as code) |
| E, a duplicate-launch alarm | 0.15% of tokens, 5 pairs a month | where-judgement-happens option 4's duplicate guard |
Prototype B and what-hides A and B are one lever: code holds the waits and says when one breaks. Built with row 3 they would be counted once.
The lines of docs/steady-base.md that change
The lines are at 066232d6. The anchor cites neither paper yet, beyond its open question and its
decision on the sixth wave, so these are the lines a fold-in of the two version 2 papers would
change. This check did not edit the anchor.
| Line | Now | Should read |
|---|---|---|
| 3 | "…survey-check-9.md checked cost-mix.md and how-process-changes-land.md. Each is now at version 2, and its figures are given here as corrected.)" |
…The sixth wave's prototype-concepts.md and what-hides-between-sessions.md were checked by survey-check-10.md, and each is now at version 2.) |
| 96 | "18 of the 25 supervisor failures on record are on a wake path, and 11 of the 15 lost wakes were lost in the runner's delivery, not in Harbour's minting." | The line stands; add: the record is a fraction of what happens. The wait graph found 27 lost wakes in September alone, 2 rescued by nothing, and the runner logs a failed [done] post as posted (22 of 23 since July) (what-hides-between-sessions.md v2) |
| 151 | "…about 8% less at list-price ratios, because the September mid tier is weighted at 0.6 of the frontier tier and listed at 0.4…" | …and the weighted unit charges both frontier price rows at 1, though per weighted token the newer row, the fleet's from the week of 22 September, lists at about 0.55 of the older; a list-price series falls there where the weighted one does not (survey-check-10.md). The rest stands |
| 170 | Row 1's risk: "Missed or lost wakes are the most common supervisor failure (15 of 25)…" | …(15 of 25 on record, and under-recorded: 27 lost in September by the wait graph, and failed [done] posts logged as posted, what-hides-between-sessions.md v2)… The rest stands |
| 172 | Row 3's evidence: "77% of supervision tokens mechanical…" | Add: code detecting broken waits caught 7 of 19 incidents on record a median 5 hours before they were filed (D1 and D2, what-hides-between-sessions.md v2); a narrow model call inside code-held steps went from 3/20 to 19/20 on a small model (prototype-concepts.md v2). The liveness clocks and a live detector are one piece of work |
| 183 | Row 14's evidence: "…one pointer probe, 51 → 18 turns (LIN-2115)" | …; John's prototypes agree and add no size: Pith's code map lost 14 of 15 to exploration, and a maintained wiki did not improve as it grew (prototype-concepts.md v2) |
| 185 | Row 16's risk: "The handoff must carry verdicts, plan text and code…" | …verdicts, plan text, code, and a list of what is done and answered: a planner given only a durable record re-proposed done work for about half its nodes (tag-two, prototype-concepts.md v2)… The rest stands |
| 190 | "From the fifth wave: 13–16. … The six papers' Options sections hold more candidates…" | …The fifth and sixth waves' Options sections hold more candidates, and some overlap: prototype-concepts' A–C are rows 14, 3 and 16, and what-hides-between-sessions' A and B sit with row 3 (survey-check-10.md); the synthesis menu will merge them and count overlapping savings once |
| 207 | "Failures that no single session can see (a dead credential healed every 30 seconds for 11 hours, LIN-3181; a circular wait that idled a passage leg for 35 minutes, 1 October) suggest a principle: code detects, the model decides. The sixth wave sizes it." | …The sixth wave sized it: 107 incidents on record since June, 92 invisible to one session, a median 16 hours to discovery (3 where the onset is timed); six rules over held data caught 9 of 19 scored (two rules 7), a median 5 hours early, and raised alarms on about 40 real faults that are on no ticket. The cost is hours, not tokens: under 2% of tokens outside quiet wakes, and 42 hours of supervisors waiting on lost wakes in September (what-hides-between-sessions.md v2). Option A is the principle at its smallest |
| 208 | "Every paper cited here has been checked by a second document. survey-check-9.md, survey-check-8.md and survey-check-7.md checked the fifth wave's six papers." |
…survey-check-10.md checked the sixth wave's prototype-concepts.md and what-hides-between-sessions.md |
| 245 | (the last evidence row, survey-check-9) |
add three rows: prototype-concepts (v2), John's prototype concepts against the levers, and what a Lighthouse study costs and catches beside a Harbour ticket; what-hides-between-sessions (v2), the failures no single session sees, and what detectors over held data would have caught; survey-check-10, the independent check of the two papers above, and every figure it changed |
Unchanged and confirmed:
- Line 207's "for 11 hours" is the dead credential's life until it was cleared. The paper's "10 hours" is from onset to filing. Both are right. Its "35 minutes" is the register's figure for the cycle's idle; the wait coding gives 62 minutes from the waiter's last turn. The two measure from different starts, and this check does not settle which the anchor should carry.
- Line 261's "a dead Linear credential re-selected every 30 seconds and reported as transient" matches LIN-3181's root-cause comment.
- Line 117's "the survey papers 10%" is what option D of prototype-concepts cites, correctly.
Method
Both papers' scripts were re-run unchanged from the repository root of this branch. Each paper
had its author's data/ snapshot copied in from the session workspace that made it, and the
outputs were diffed against those snapshots. Nothing was written to the authors' copies.
# what-hides-between-sessions, from the author's snapshot (data/survey-hides)
node scripts/survey-hides-analyse.mjs --dir data/survey-hides [--near]
node scripts/survey-hides-detect-sessions.mjs --in data/survey-hides
node scripts/survey-hides-code-waits.mjs --in data/survey-hides
node scripts/survey-hides-detect-runner.mjs --wake data/survey-hides/wake --out <dir>
node scripts/survey-hides-figures.mjs --in data/survey-hides/analysis.json --out <scratch>
# and fresh, into data/survey-hides-rerun: transcripts --since 2026-08-29, runner,
# survey-wake-extract.mjs --since 2026-08-29 --out <dir>/wake, detect-runner, detect-sessions,
# code-waits, analyse; then the same with sessions cut to a first turn before 2026-10-01T10:14Z
# prototype-concepts, from the author's snapshot (data/survey-protoconcepts)
node scripts/survey-protoconcepts-lighthouse.mjs
node scripts/survey-protoconcepts-harbour.mjs
node scripts/survey-protoconcepts-figures.mjs --out <scratch>
node scripts/survey-costmix-tokens.mjs --since 2026-09-01 --until 2026-10-01T10:00:00Z --out <dir>/tokens.json
node scripts/survey-protoconcepts-dollars.mjs --out <dir>/dollars.json
# this check: price rows (the dollars census, each message also tallied by row, session and week)
node scripts/survey-check-10-rows.mjs --out data/survey-check-10/mix.json
node scripts/survey-check-10-rows-analyse.cjs data/survey-check-10
- Weighted tokens for Lighthouse are estimates. Each study's frontier-row spend is divided by the median per weighted token of its 9 frontier dispatches with token records, and its top-tier review spend by that of its 3 review records.
- The prototypes were read with
git show <sha>:<path>at the matrix's pins, so no checkout state matters. - The checks coding. Every third (check, document) pair, 8 in all, was coded from the check's own text before the JSON was opened. The rest were read where a check draws its own line between load-bearing and secondary.
- The alarm sample. The definitions were written before any alarm was read. Each sampled alarm
was read in the transcripts of the sessions it names and in the runner's oplog around it. The
codes, with evidence lines, are in
survey-check-10-alarm-codes.json. - The overlap with row 1 joins the 270 re-fire deliveries of
data/survey-hides/waketo the layer of the session each entered. A quiet one is an outcome of none or read, and a supervisor is an autopilot, leg, stepper or Runner session. - Proxy. Reads of the ticket and its brief, and nothing else. The only writes are this ticket's comment and status.
Limits
- Every reader shares a tier with the authors. The re-runs and readings are independent of the authors' sessions, not of their model tier.
- The data are the authors' snapshots. A re-run proves that the scripts are deterministic over them, not that the snapshots were complete. The fresh extractions say the transcript and oplog snapshots were; the tracker snapshot was not re-taken.
- The alarm sample is 33 of 93, coded by one reader with fixed definitions. The interval on "about 40" is wide, 25–50, and the line between a real fault and a harmless blip is a judgement. Bias: a reader looking for faults may code a blip real, which runs the count high. Usage-limit stops count as real here; without them it is about 33.
- "Load-bearing" by the checks' own word depends on how carefully each check chose its words. The coding's count and this one bound the truth, at 23 and about 20.
- The price rows are Lighthouse's table as read on 26 September. List prices that changed within September would move the per-row figures. The finding that a population's price per weighted token follows its row, not its shape, holds at any prices.
- Lighthouse's output estimate is sized on Harbour's transcripts, not its own.
- The anchor may move. The lines are listed at
066232d6; a later edit shifts them.
Next
- Does the scorecard's weighted unit hide September's change of frontier price row? The weighted
unit charges both frontier rows at 1, but per weighted token the newer row lists at about 0.55 of
the older, and the fleet moved to it in the week of 22 September. Re-price September's cost per
correct change by week at each session's own row, and say how much of any fall after 22 September
the weighted series would credit to the process when it was the price. This goes into
proposals.md. - Would D1 and D2 run live be worth reading? The detector proposal already in
proposals.mdnow carries this check's figures: D1 needs a duration or harm threshold, since its unrecorded bursts were nearly all harmless, and D2 needs its three-minute merge fixed and its completed waiters set apart before a shadow run's alarms are counted.