---
title: Does the evidence support the agent-retention paper’s conclusions?
kind: check
version: 1
date: 2026-09-20
authors: [Codex reviewer]
model: "Codex in ChatGPT Work; separately authored by the paper_checker subagent. Exact model identifier and effort not exposed. Same model family as the main author; separate checking pass, not external replication or Harbour dispatch lineage."
grounded_at: "LinearViewer 4b28d25c234b011ccd7c6fa795645e13598623e6; external snapshots and reviewed draft SHA256 values below"
cites:
  - "what-should-an-agent-leave-behind.md, SHA256 f1e736ac98f239528a9c2eda433a2e067ab10483b907171188f8ff9619908a17"
  - "what-should-an-agent-leave-behind-evidence.md, SHA256 8c7785e974a3b3fbe358e4daf464e581a110dce6b13ae5493c9c10d59ae6513d"
  - "tangle/PROGRESS.md@535b96a20cbcf9303f0955c02139dbf7777fd179:166-193"
  - "docs/reviews/context-efficiency-ceiling-review-2026-08-15.md@4b28d25c234b011ccd7c6fa795645e13598623e6:196-251"
  - "docs/papers/harbour/review-consumption-marks.json@4b28d25c234b011ccd7c6fa795645e13598623e6"
---

# Does the evidence support the agent-retention paper’s conclusions?

Yes, within its stated limits. The [paper](what-should-an-agent-leave-behind.md) supports testing retained artifacts against later use and total cost; it does not establish that wikis save work, graphs generally outperform ordinary tools, or retained records reduce human supervision. I found no blocking factual or inference problem in the checked central claims. Independent arithmetic reproduced the published comparisons, but this pass did not independently establish the truth of their original outcome labels.

## Findings

**The Tangle account includes its counterexamples and positive reference-model results.** I checked all 18 composing-control rows against the [published comparison](https://github.com/JKershaw/tangle/blob/535b96a20cbcf9303f0955c02139dbf7777fd179/PROGRESS.md#L166-L193), loaded each raw control and its exact `limits.matched` graph, and recomputed totals from the result entries. Every appendix count, token sum, run-commit field and dirty-state qualifier matched. All pairs record `repeats=1`. Graph coverage is higher in 14 rows, equal in one and lower in three. Both API models favour the graph on both evaluated suites; the 8B MangoDB row favours the control. This confirms the appendix’s selection and arithmetic, not a statistically independent set of 18 successes or failures. The main paper’s four MangoDB rows faithfully illustrate that range.

**Tangle’s endpoint cannot establish semantic correctness or a topology effect.** The [grader](https://github.com/JKershaw/tangle/blob/535b96a20cbcf9303f0955c02139dbf7777fd179/scripts/grade.js#L59-L87) checks keyword alternatives in findings and cited records. The [composing control](https://github.com/JKershaw/tangle/blob/535b96a20cbcf9303f0955c02139dbf7777fd179/scripts/tools.mjs#L1-L13) treats something named and read as supported, while the cited variant restricts output to copied sentences. Those are different output constraints and support conventions. The paper correctly declines to treat support as entailment or to attribute the joint intervention in extraction, sequencing, selection and stopping to graph structure alone. Its token comparisons also avoid asserting cross-model price equivalence.

**The pointer result survives arithmetic checking and remains a hindsight probe.** The [source table](https://github.com/JKershaw/LinearViewer/blob/4b28d25c234b011ccd7c6fa795645e13598623e6/docs/reviews/context-efficiency-ceiling-review-2026-08-15.md#L196-L251) gives pointer-run costs $0.357/$0.491/$0.334 and turns 16/30/18: medians $0.357 and 18. Against the one no-pointer run, $0.783 and 51 turns, the ratios are 2.19 and 2.83. All four are reported to pass 43 checks. The paper preserves the decisive qualifications from the [report’s limits](https://github.com/JKershaw/LinearViewer/blob/4b28d25c234b011ccd7c6fa795645e13598623e6/docs/reviews/context-efficiency-ceiling-review-2026-08-15.md#L610-L641): hindsight, excluded acquisition cost, one no-pointer run and a behaviour verifier narrower than the deliverable. The 93-token difference is an estimate from bytes/4. This supports the existence of a useful handoff in one recorded case, not automated pointer construction or net lifecycle savings.

**The consumption figures are correct and describe an instructed consumer.** Running the [existing recomputation](https://github.com/JKershaw/LinearViewer/blob/4b28d25c234b011ccd7c6fa795645e13598623e6/scripts/review-consumption-recompute.mjs#L1), then separately expanding and intersecting the committed annotation ranges, reproduced 181/187 ledger units, 109/429 method units and 463/957 units overall. The independent calculation also checked section ranges for overlap and consumed identifiers for membership. The [close-out template](https://github.com/JKershaw/LinearViewer/blob/4b28d25c234b011ccd7c6fa795645e13598623e6/lib/prompt-template-defs.js#L1113-L1127) expressly requires ledger consumption. The paper therefore has grounds for its consumer-specific design inference, while correctly withholding claims about human attention, resolution quality or the consequences of removing other prose.

**The remaining historical comparisons support caution about proxies, not causal verdicts.** Pith’s [15 score tables](https://github.com/JKershaw/pith/blob/d52c0151917ece0188a6f941f48df4edb400485d/docs/benchmark-results/2026-01-03-self-test.md#L30-L308) independently sum to 236 and 286 of 300 rubric points, with zero Pith wins, one tie and 14 losses. Its [judge definition](https://github.com/JKershaw/pith/blob/d52c0151917ece0188a6f941f48df4edb400485d/docs/BENCHMARKING.md#L284-L320) supplies the control as ground truth. CodeWiki’s [checkpoint report](https://github.com/JKershaw/CodeWiki/blob/154aa877d9ee2c32a4b6515ff10b6c4d9b59a84b/docs/wiki-quality-analysis-2025-12-09.md#L1-L38) agrees with the page and score sequences. TAG-two’s five named task files contain 6, 12, 19, 4 and 10 operations, totalling 51; its [planner account](https://github.com/JKershaw/tag-two/blob/51476740532a1b18ab60aaedac38200a9d40dfd6/docs/experiments.md#L1162-L1187) itself identifies non-entailing evidence in its labels. Browser-agent’s [chain log](https://github.com/JKershaw/Browser-agent/blob/68d7f4c263d8ca20613bd2ed231035d3ed343678/BUILD_LOG.md#L1210-L1264) and [task definition](https://github.com/JKershaw/Browser-agent/blob/68d7f4c263d8ca20613bd2ed231035d3ed343678/scripts/eval/tasks.js#L135-L149) support a bounded second-hop dispatch result with a shared transcript. The paper’s distinctions between these endpoints and completed downstream work are warranted.

## Method

This was a separately authored source and arithmetic check on 20 September 2026. I read the standard, main draft and appendix, confirmed the six repository HEADs against their cited snapshots, read the cited historical passages and relevant grader/template code, and performed the deterministic calculations described above. For Tangle, I also independently enumerated composing `mode=tools` result files at the three published control revisions (`a811bbd`, `f2722b8`, `4aa0cdd`): the resulting 18 files exactly matched the appendix, with no missing or extra files. Earlier development variants and cited controls were outside that stated population.

The reviewed main draft’s SHA256 was `f1e736ac98f239528a9c2eda433a2e067ab10483b907171188f8ff9619908a17`; the appendix’s was `8c7785e974a3b3fbe358e4daf464e581a110dce6b13ae5493c9c10d59ae6513d`. These identify the uncommitted text actually checked. The external snapshots are those in the linked sources; Harbour sources were read at `4b28d25c234b011ccd7c6fa795645e13598623e6`. No model trials were run.

## Limits

I did not relabel Tangle facts, Pith answers, TAG-two repetitions or consumption units. I did not recover unavailable raw Browser-agent trials or Harbour probe sessions, re-audit bills, reconstruct dirty historical working trees, or read the live private tracker records. Harbour’s record-size findings were checked against the cited paper, not rederived from its underlying ticket population. I did not independently audit the 39-repository inventory or the prior-work literature review. This is a separate pass within the same Codex model family, informed by the draft’s selected sources; it is neither external replication nor an independent sampling exercise. Correct arithmetic cannot remove biased labels, retrospective selection or missing outcomes.

## Next

For the proposed caller/source/index/explanation comparison, freeze a small unseen task set and its acceptance rubric before constructing retained material. Have a separate checker adjudicate task outcomes from source and execution evidence without seeing which arm produced them; report disagreements and failed runs alongside total construction, retrieval, inference, verification and maintenance cost. Add this adjudication requirement to the existing LIN-2689 follow-up rather than creating another claim from the current scores.
