Evidence register: what should an agent leave behind?

Supporting appendix to the paper, version 1, 20 September 2026, tracked as LIN-2961. This register distinguishes recorded historical results, arithmetic recomputation, selected source checks and proposed interpretations. It is not a new effectiveness experiment. The paper and this register were assembled by Codex with separate source-audit contributions.

Scope and selection

The initial public inventory contained 39 repositories. Sixteen central/supporting repositories were inspected at pinned snapshots, seven adjacent projects received lighter notes, and three relevant-looking repositories had no readable contents. Those relevance notes provided discovery, not a statistical sampling frame. This paper selects five experiment-bearing projects (CodeWiki, Pith, tag-two, Tangle and Browser-agent) and Harbour records directly addressing retention, handoffs or consumption. TAG, CodeWiki-Generator and the other implementations supply background rather than additional outcome estimates.

Harbour documents are pinned to 4b28d25c234b011ccd7c6fa795645e13598623e6. The five external repository snapshots appear below. Historical run versions are reported separately, including missing or dirty run states. Selection is retrospective and favours projects with readable evaluation records. The claim table includes favourable and unfavourable observations; the Tangle table includes all 18 composing-control comparisons in the inspected progress table, rather than selecting only wins or losses. No pooled score across projects is meaningful.

The live Harbour reads on 20 September 2026 were GET /issues/LIN-2911, /issues/LIN-2442, /issues/LIN-2716, /issues/LIN-2115 and /issues/LIN-2689 through the authenticated workspace proxy. They corroborate selected paper records and the existing follow-up task; they are not a re-audit of the full underlying cohorts. Raw private response bodies and credentials are not included in this publication. Public source links and exact comment identifiers below let an authorised reader repeat those checks.

Cross-project evidence

ID Defensible claim Endpoint and denominator Run state versus inspected snapshot Counterevidence and causal limits
C1 More generated pages did not improve the measured quality of one historical CodeWiki run. Three reported checkpoints after 150 iterations: 30→53→62 pages; quality 59.57→58.11→58.00; fully accurate answers 6/15→3/15→4/15. Endpoint is internal wiki/answer grading, not caller productivity. Report dated 2025-12-09; model meta-llama/llama-4-maverick. Run commit not recorded in the report; inspected commit 154aa877d9ee2c32a4b6515ff10b6c4d9b59a84b. Growth accompanied parsing and linking failures. This does not isolate content volume as a cause, prove a statistically reliable decline, or describe later code. Later verification/consolidation code cannot retroactively repair this result.
C2 Pith returned lower-scoring answers than direct exploration in its recorded self-test; that is not an estimate of caller savings. 15 tasks × four 1–5 criteria: recomputed totals 236/300=78.6667% and 286/300=95.3333%; difference −16.6667 percentage points; 0 wins/14 losses/1 tie. These are rubric points, not success percentages. Run report says version 7c36b5f, 2026-01-03, qwen/qwen-turbo; inspected d52c0151917ece0188a6f941f48df4edb400485d. Comparator/judge model identity and resource parity are not adequately established by the report. The query judge calls the control answer GROUND TRUTH. Some answers were truncated. One task ties perfectly; scoring assigns all control correctness/completeness/specificity 5/5. Self-testing and tuning limit generalisation; no caller-with/without-Pith trial.
C3 Durable task records supported inspectable completion but did not reliably prevent repetition or unsupported inherited claims. Five archived task files independently counted: 4+6+10+12+19=51 operations. Two external-repository episodes contribute 31 operations. Records/reports show one prevented repeated refused close, four repeated passing verifications and an empty-stdout fabrication. The episodes evolved the implementation; no single frozen run commit identified here. Inspected 51476740532a1b18ab60aaedac38200a9d40dfd6. Each raw file supplies operation/state records, not a complete exact executable environment. Positive evidence matters: reports say three closed episodes were correctly reconstructed by stateless reader calls; two Harbour episodes reached verified completion and recovered from several issues. No no-record control, no two-real-process agent handoff, and steward-selected tasks; a zero exit does not establish appropriate completion.
C4 A planner could repropose work already answered while receiving the durable record. Repository check reports 24/47 proposed nodes across 11 supplied-record graphs (51.06%), versus 4/14 across 3 same-era comparison graphs (28.57%). Node labels are model/check judgements, not independently audited truth. Historical evolving planner runs, not a single commit. Inspected 51476740532a1b18ab60aaedac38200a9d40dfd6. Source account maps checks against outcomes available at each run. Not a causal estimate that memory worsens planning: controls used another commit, only 3 graphs, check varies around one node per graph, lexical overlap could bias matching. Two of three done verdicts in one run cited non-entailing evidence. This counterevidence prevents treating the 24 labels as verified facts.
C5 A deterministic chain driver improved a bounded tool-dispatch task in the recorded Browser-agent trials. Second-hop task: 3/20 before; 19/20 then 20/20 after (39/40 pooled=97.5%). Qwen3-0.6B, temperature 0.6. Endpoint checks correct Wikipedia request and answer pattern, not general task solving. Reported wall time 3.6→4.1 seconds. Run commit not identified in the relevant log section. Inspected 68d7f4c263d8ca20613bd2ed231035d3ed343678. Task and scorer code retained; baseline/after raw sample files not found among tracked files during this audit. Before/after development result, not independent replication across tasks. Shared transcript retained; intervention is sequencing and explicit turns, not proof of context splitting. Pooling regression/dev samples does not create an untouched confirmatory trial. Harder recall/fan-out tasks remained weak.
C6 Tangle had mixed coverage–cost tradeoffs; the full comparison does not justify a broadly negative graph verdict. All 18 published composing-tool-control comparisons reconstructed below from raw JSON and each control file’s matched pointer. 42 MangoDB topics across 4 briefs, 35 profile topics across 4 briefs, 30 overflow topics across 3 briefs, 23 question facts across 7 seeds. Single archived run per configuration (repeats=1). Run commit strings and raw graph/control artifacts appear in full table; inspected 535b96a20cbcf9303f0955c02139dbf7777fd179. Several run commit fields explicitly say uncommitted changes. The 8B MangoDB reversal is real; both API models show higher graph coverage on profiles and MangoDB. More coverage often consumes much more input/output tokens. Selection, sequencing, deterministic extraction, graph storage and stopping differ together; the comparison does not identify a topology effect or a general productivity effect.
C7 Source provenance and semantic support must be distinguished in interpreting Tangle scores. gradeRun marks a fact present via keyword alternatives in the finding; supported additionally requires a cited record containing matching alternatives. Graph extraction copies source, whereas composing control support means answer mentions a fact found in what was read. Current grader inspected at 535b96a20cbcf9303f0955c02139dbf7777fd179; historical run metadata ties outcomes to prompts/schema/commits. Full historical scorer parity not exhaustively audited. This is not a semantic entailment, relevance or contradiction check. Reported support labels can overstate epistemic certainty; extraction of irrelevant comments/examples remains possible. The near-zero cited-control column imposes a different output constraint and should not be represented as general tool-agent accuracy.

Claim sources and raw-material status

C1 — CodeWiki

Dated report and table; Recorded implementation failures. The audit checked these recorded numbers, not the original judge outputs. Baseline and later checkpoints are not independent replications. Do not equate the later inspected commit with the experiment version.

C2 — Pith

Run configuration and pipeline; Fifteen task score tables; Reported aggregates; Judge definition: control labelled ground truth. Arithmetic recomputed by summing each of the fifteen bold Total rows. Full raw model/judge transcripts were not identified for this report; displayed scores and narrative are the accessible evidence. The 99-second/$0.30–0.50 construction entries are reported/estimated, not independently billed economics.

C3–C4 — TAG-two

Limitations and summary numbers; Planner comparison and caveats; Three task-runner episodes and positive resumption findings; Two Harbour episodes and failures.

Raw task states (operation counts independently recomputed):

The operation count is exact for these snapshots. The classification of repeated work and fabricated conclusions remains an attributed interpretation of the records; it was not independently re-labelled here. Planner graph examples and historical check discussion are retained, including run87 checked graph and run88 graph. Their existence does not independently validate all 24/47 labels.

C5 — Browser-agent

Chain-driver mechanism, rates, runtime and model; Actual task endpoint; Scorer; Holdout checkpoint and task scope; Task ladder and limitations. The log names holdout-after-chain.json, but that raw file was not found in tracked files. test-results/ is ignored. Recover raw trials before calling the percentages independently audited. Do not confuse the separate 29/40→39/40 URL-hint result earlier in BUILD_LOG with this chain result.

C6–C7 — Tangle

Complete published control table; Reference-model interpretation and limitations; Actual keyword-based grading; Meaning of composing/cited control support.

The next table includes every composing-control row in the published Sept 19 table. Earlier tools-1/tools-2 development variants are excluded consistently; modes ending in a811bbd, f2722b8 or 4aa0cdd are the published tools-3/follow-on rows. Counts come from raw JSON totals; tokens are the sum of each result.cost.tokens; graph selection follows limits.matched, not a guessed filename. Each listed pair reports repeats=1. Dirty denotes the raw commit field containing “(uncommitted changes)”. The archive pins outputs precisely, but dirty runs require additional source/prompt state for exact reproduction.

Suite / model Graph present / total; tokens Control present (supported) / total; tokens Graph run commit Control run commit Raw pair
seeds / qwen3:1.7b 15/23; 15,555 14 (0)/23; 17,625 0b0ed70 + dirty a811bbd graph · control
seeds / qwen3:4b 12/23; 24,371 18 (0)/23; 29,439 877f185 a811bbd graph · control
seeds / qwen3:8b 12/23; 19,038 17 (15)/23; 18,435 877f185 + dirty a811bbd graph · control
seeds-code / anthropic/claude-haiku-4.5 25/42; 128,085 20 (20)/42; 47,622 4aa0cdd + dirty 4aa0cdd + dirty graph · control
seeds-code / deepseek/deepseek-v3.2 13/42; 67,440 8 (7)/42; 75,547 054ad09 4aa0cdd + dirty graph · control
seeds-code / qwen3:1.7b 14/42; 40,391 6 (6)/42; 36,592 0b0ed70 + dirty a811bbd graph · control
seeds-code / qwen3:14b 18/42; 62,340 15 (15)/42; 32,241 a811bbd f2722b8 graph · control
seeds-code / qwen3:4b 11/42; 47,339 11 (6)/42; 43,482 877f185 + dirty a811bbd graph · control
seeds-code / qwen3:8b 12/42; 53,179 15 (14)/42; 15,426 59c9099 a811bbd graph · control
seeds-overflow / qwen3:1.7b 24/30; 43,855 11 (10)/30; 38,503 0b0ed70 + dirty a811bbd graph · control
seeds-overflow / qwen3:4b 23/30; 49,522 11 (10)/30; 34,786 877f185 + dirty a811bbd graph · control
seeds-overflow / qwen3:8b 24/30; 51,468 21 (20)/30; 18,056 59c9099 a811bbd graph · control
seeds-profile / anthropic/claude-haiku-4.5 23/35; 138,757 15 (15)/35; 45,612 4aa0cdd + dirty 4aa0cdd + dirty graph · control
seeds-profile / deepseek/deepseek-v3.2 27/35; 114,524 21 (20)/35; 105,862 4aa0cdd + dirty 4aa0cdd + dirty graph · control
seeds-profile / qwen3:1.7b 25/35; 70,770 13 (13)/35; 49,143 0b0ed70 + dirty a811bbd graph · control
seeds-profile / qwen3:14b 27/35; 87,481 20 (18)/35; 15,549 f2722b8 f2722b8 graph · control
seeds-profile / qwen3:4b 24/35; 75,855 11 (10)/35; 63,653 877f185 a811bbd graph · control
seeds-profile / qwen3:8b 24/35; 85,043 20 (17)/35; 25,142 59c9099 a811bbd graph · control

Interpretation: the graph has more present rubric items in 14 of these 18 comparisons, ties in one, and fewer in three (the 4B/8B familiar-question suite and 8B MangoDB). This row count is descriptive, not a statistical meta-result: suites/models share tasks and development history, endpoint meanings vary, and configurations are not independent samples. The previously emphasised 8B code case must be shown alongside the positive API-model rows. The graph does not uniformly cost fewer tokens; comparing input/output sums is not identical to billed compute or latency.

The support qualifier matters: familiar-question control answers can score high from model memory while recorded source-support is zero. Conversely, source keyword matches do not verify reasoning. A cited-control arm demands verbatim extraction and performs differently; include it if discussing extraction constraints, but do not silently substitute it for the ordinary composing control.

H1: review consumption

Claim. An explicit downstream contract is associated with visibly different consumption across sections of an agent's record. This is a consumption result, not a causal productivity result.

Section Annotated units Marked consumed Share
Ledger 187 181 96.8%
Method/check narration 429 109 25.4%
All sections 957 463 48.4%

Evidence and checks. The version-2 paper, committed annotations and recompute script were read. Running node scripts/review-consumption-recompute.mjs in the pinned checkout reproduced the table. This recomputes the author's annotations; it does not independently validate their meaning.

Two cited review/close-out pairs were inspected through the live proxy: LIN-2442, review 1d7c33a7-5d93-4501-8e18-2f00d57da73c (1 September) and close-out 37e3ebcd-34cf-43a1-9804-e64af6c3aaa6 (2 September); LIN-2716, review prefix 3a77e67a and close-out prefix 3b08385b (11 September). They corroborate explicit ledger disposition, not every sentence label. LIN-2911 review 3106897e-68e2-4446-aa2a-c614fa2c3a09 (18 September) documents the numerical faults in version 1 that version 2 corrected. Our synthesis uses the corrected version.

Limits. The ten reviews were selected under one template and an eighteen-day window. Re-reading a check counts as consumption and can inflate the apparent transfer of useful information. The annotations count units, not reading time or words. Unobserved human readers and other agent consumers are outside the study. The templates specifically direct close-out to consume the ledger, so higher ledger consumption is partly an enacted interface contract. Disposition can mean routing a problem elsewhere, which does not demonstrate its resolution. None of these figures estimates the effect of removing unconsumed material.

H2: pointer handoff

Claim. A historical same-model probe provides an example of a retained source pointer reducing subsequent search effort; it does not establish automated knowledge creation or lifecycle savings.

LIN-2078 probe condition Runs Median/report cost Median/report turns Behaviour verifier
Requirement plus source site, Haiku 4.5 3 $0.357 18 43/43 in each run
Requirement without source site, Haiku 4.5 1 $0.783 51 43/43

Evidence and checks. The context-efficiency ceiling report, sections 4.2–4.3 gives individual pointer-run costs $0.357/$0.491/$0.334 and turns 16/30/18. Sorting these reproduces the medians. The ratios are $0.783/$0.357 = 2.19 and 51/18 = 2.83. The reported 93-token difference is a bytes/4 estimate, not a tokenizer measurement. Probe worktrees started at 04acdb20; the landed regression file was swapped in after execution as a behaviour verifier. The inspected report is later than that run. LIN-2115 comments from 15 August confirm the report's review/correction/merge sequence; the final close-out c7f0a45e-05ca-4e3c-9c85-8671b8acce84 names merge 7643cacf0ca4a31296acde4536a249cbcdbcdb1a and PR #1143.

Limits and counterevidence. Raw sessions, prompt distillates and produced diffs are in an on-machine bundle unavailable here. This review checks the published table and corroborating tracker record, not the raw billed usage or execution. The handoff was written with hindsight, and the time/cost of obtaining its pointer is excluded. The no-pointer condition has only one run. The verifier assesses production behaviour rather than the full deliverable, including the quality of the tests written by the worker. Section 10 reports a task with no passing rung, an invalid review verifier, answer leakage caught in development, and multiple confounds in the much larger whole-leg saving claims. We do not carry those larger savings into this paper. The report distinguishes a historical transcript-pricing defect from the probe's first-party cost figures, but we have not re-audited either billing source.

Existing continuation. LIN-2689, description read on 20 September, proposes deterministic indexes and pointer-first handoffs and explicitly references CodeWiki-Generator and Pith. It remains a proposal, not evidence that the automated index has achieved the probe's gain. LIN-2961 is linked to it as related work; this paper does not claim the proposal was newly discovered or implemented here.

H3: record size and outcomes

Claim. An observational Harbour study did not find improved review outcomes with longer records in its cohort; it does not identify the causal effect of shortening records.

The version-2 paper reports 184 Done tickets and a reviewed subset of 111. Its thirds of pre-review writing had first-pass approval rates of 66%, 68% and 64%. This review read the paper and its limitations, not a fresh export and reclassification of all tickets. These remain attributed historical estimates.

Task size, task difficulty, workflow selection, review standards and the short observation window can affect both record length and measured outcomes. Short lanes preferentially handle small work. A later defect count depends on detection and classification; the paper was itself revised after a checking paper distinguished review residue from genuinely later-found faults. Neither equal observed approval nor an absence of detected defects establishes equivalent real quality. We therefore use this study to justify measurement and caution about verbosity as a proxy, not to prescribe deleting particular sections.

Four mechanism-level precedents were selected and their relevant full-text method/experiment sections read. This is a bounded contextual review, not a systematic or comprehensive literature search.

Reference Retained object or process Why the present records are not a replication
Packer et al., MemGPT, v2, 12 February 2024, sections 2–3 External conversation/document memory accessed through explicit functions Conversation baselines receive compressed history; memory access and retrieval opportunities differ. No matched coding-wiki economics are tested there.
Shinn et al., Reflexion, v4, 10 October 2023, sections 3 and 4.3 Feedback-derived lessons across task trials Actor, evaluator and reflection generation jointly determine outcomes. TAG-two's stored task claims and evaluation conditions differ.
Wang et al., Voyager, v2, 19 October 2023, sections 2.2–2.3 and 3 Executable skills retrieved and composed during later Minecraft tasks Executable skills, curriculum and environment feedback differ from repository explanations and software-maintenance tasks.
Besta et al., Graph of Thoughts, v4, 6 February 2024, sections 4.4–4.5 and 5 Generation, aggregation and evaluation under a user-constructed operation graph Tangle's investigation policy is not this protocol; tasks, scoring, budgets and graph construction differ.

These precedents rule out presenting persistence or structured reasoning as new ideas. They do not supply directly comparable benchmark numbers. This paper makes no claim to replicate, refute or supersede them.

Reproduction and interpretation checklist

  1. Check out the cited repository snapshots. Read run-version fields separately from the snapshot that archives them. Do not silently fill missing versions.
  2. For Pith, sum the fifteen displayed Total rows (236 and 286) against 300 possible rubric points. For TAG-two, count the five raw task files' operations (6, 12, 19, 4 and 10).
  3. For each of the 18 Tangle control JSON files linked above, read limits.matched and load that exact graph file. Read factsPresent, factsSupported and factsTotal; sum each entry's cost.tokens across results. Preserve commit, repeats, prompt/schema and runtime metadata. Do not deduce the matched graph from the model name.
  4. Run Harbour's existing consumption-recompute script. It verifies counts against saved labels, not label validity. For live corroboration, an authorised reader can retrieve the five named issues and locate the cited comments.
  5. Treat each endpoint as its own measurement. Do not turn rubric points into task-success percentages, keyword support into entailment, consumption into human attention, or probe cost into total lifecycle cost.

No new model trials, deployment changes or paid benchmark calls were performed for this retrospective. The separate checking paper records which of these checks another author repeated and what it could not verify.