---
title: The Second Copy
kind: essay
argument: A decision held in more places than any one change reaches costs nothing while it sits and is paid in full when it changes; fifty years of research puts the cost exactly there, the new agent studies show agents make more such places and delete fewer, and Harbour's record shows the bill arriving at the review gate and in the later sweep rather than in the hours; so coherence is kept by structure that leaves one place to change and by making deletion the normal end of a change, not by rules or reviews that ask for it.
version: 1
date: 2026-10-08
authors: [Claude, for John Kershaw]
model: "claude-fable-5-1, Claude Code on the web, the same interactive cloud session that wrote coherence-as-it-grows.md (versions 1 to 3) and ran its blind readers; no dispatch lineage. No new measurement: every Harbour figure is read from that paper or the archive papers at the commits named in the sources, and every outside figure from the primary source the paper's literature check read. A tone and readability review and an accuracy review by two in-session subagents preceded this version, and their thirteen corrections and ten edits are in it."
sources:
  - "docs/papers/harbour/coherence-as-it-grows.md@95b317fb (version 3): findings 1 to 7, Options, How to test it, Limits"
  - "docs/papers/harbour/coherence-as-it-grows-check.md@95b317fb (version 1, the independent check of the paper's version 2)"
  - "coherence-as-it-grows-sources/lineage.md@95b317fb and coherence-as-it-grows-lineage.json@95b317fb (the four decisions' lineage)"
  - "docs/papers/harbour/coherence-as-it-grows-sources/check.md@95b317fb (fifteen sources read at the primary)"
  - "docs/papers/harbour/measuring-throughput.md@d61903f3:19-20; why-throughput-halved.md@d61903f3:19-23; reliability-baseline.md@d61903f3:16-17,48,51; steady-base-menu.md@d61903f3:64 (the decline and its attribution, read at those lines)"
  - "CLAUDE.md@a82b8758 (the first commit, 2026-01-04), line 36, the 'no build step' line; public/swim.js@d61903f3:29-31"
  - "Outside sources: each entry in the annotated reading list carries its DOI or URL; the markers in the text are the reading list's."
---

# The Second Copy

*What fifty years of research, the new agent studies, and one codebase built by agents say about keeping a system whole*

Prepared for John Kershaw | October 2026

Every duplicate in a codebase was cheaper to make than the alternative, on the day it was made. That is why there are so many, and why asking people not to make them has never held for long. The research that bears on this is older than most of the people who argue about it, and it has a settled answer that nobody likes: the duplicate costs nothing while it sits there. It costs when the thing it duplicates has to change. Then the change lands in fewer places than it needed to, and a fault turns up later, somewhere nobody was looking. Agents, it turns out, make more duplicates than people and delete fewer, for reasons psychology had already seen in people. And a codebase built almost entirely by agents, over nine months, shows the bill arriving exactly where the old research said it would: not in the hours, which are hard to see, but in the review that sends the change back and in the sweep that finds the copy the fix never reached. What follows is the argument that these three bodies of evidence are one story, and what the story says about keeping a system whole. The short version is that you cannot ask for coherence. You have to build the one place to change, and you have to make deleting the old one part of what finished means.

## 1 A line in the first commit

Harbour's first commit, on 4 January 2026, carried a rules file with one line that would turn out to matter more than any other: `Keep it minimal - no frameworks, no build step`. It is a good rule. It is also, read a certain way, a rule that the browser and the server may never share a module, because without a build step the usual way of sharing one does not exist. Nobody wrote that second reading down. It was simply what each agent concluded, independently, when it needed a function on the browser side that already existed on the server side. So the function was copied, and a comment was left. The one in the swim-lane view reads: "duplicated because the `no build step` constraint (CLAUDE.md) prevents importing from lib/. Keep in sync; unifying client+server is a candidate follow-up under LIN-174." It has pointed at that ticket since 7 June. The ticket closed on 10 June with no such follow-up filed under it.

By October the codebase held thirty-three decisions whose comments admit a second copy kept in sync by hand, and a blind second reading found twenty-five of them still in the code. A tolerant timestamp parser, the kind that returns a number or nothing rather than throwing, existed in one file in May and eighteen in October. The bag of options a route passes to the page shell went from one file to fourteen. Not one of the eighteen timestamp parsers carries a comment naming another. They were not copied. Each was re-derived, by an agent that needed one and wrote one. [H1]

That is the whole phenomenon in miniature. A reasonable rule, a reasonable reading of it, a reasonable copy, and a comment that files the debt against a ticket that will never pay it. The rest of this essay is about why that pattern is so stable, what it costs, and what, if anything, interrupts it.

## 2 What Parnas was pointing at

The clearest statement of the problem is also the oldest. In 1972 David Parnas proposed that a module should be chosen to hide "a design decision which is likely to change" [P72]. The criterion is usually taught as a principle of good design. It is more useful read as a prediction. A decision that is likely to change, and is implemented in several places, is one that was never hidden; and when it changes, every place must change with it, or the ones that do not become faults. Parnas said as much seven years later, naming "excessive information distribution" as the first reason a system becomes hard to extend or, his telling word, to *contract* [P79]. His example was a decision to support three languages that had leaked into tables sized for exactly three, so that removing one would have cost as much as adding one. In 1994 he gave the process its name, "ignorant surgery": changes made without the design concept in mind, after which "a change that might have been made in one or two parts of the original program now requires altering many sections of the code," and "it is more difficult to find the routines that must be changed" [P94].

Read those three papers together and the mechanism has a shape. A decision scatters; the scattering raises the cost of the next change to that decision; and the scattering also raises the chance that the next change misses a place, because finding every place is itself work. Perry and Wolf, defining architecture in 1992, added the loop that makes it self-sustaining: drift produces "a lack of coherence and clarity of form, which in turn makes it much easier to violate the architecture" [PW92]. Brooks had already said what the goal looked like from the other side. Conceptual integrity, one set of ideas rather than "many good but independent and uncoordinated ideas," was to him the most important property of a system [B75], and he treated the removal of duplicates as part of what software *is*: "no two parts are alike. If they are, we make the two similar parts into one" [B87].

None of this was measured. It was argument from experience, right about the mechanism and silent about the size. The measuring came later, in two camps that spent twenty years apparently disagreeing.

## 3 The clone wars, and how they end

The first camp measured duplication and found it costly. Eick and colleagues had fifteen years of change records for a telephone switch, and found that the share of changes touching more than one file more than doubled across those years, and that the span of a change predicted its effort [E01]. Mockus and Weiss, on the same system, found that how far a change was spread across files, modules and subsystems was "essential to predicting failures" [MW00]. Eaddy and colleagues found that the more scattered a concern's implementation was, the more defects it had, and gave the obvious theory: "maintainers may make changes incorrectly or neglect to make changes in all the right places" [E08]. Hassan found that how a change scattered across files predicted faults better than how much code it churned [H09]. Juergens and colleagues took five systems and asked what happened when a clone was changed. Nearly every second unintentional inconsistent change was a fault, confirmed by the developers who owned the code [J09]. This essay leans on that result more than any other.

The second camp measured duplication and found it harmless. Kapser and Godfrey judged that as many as 71% of the clones in two open-source systems could be considered good for maintainability, and titled their paper to say so [KG08]. Rahman, Bird and Devanbu found cloned code no more defect-prone than the rest, and if anything less, with clones spread across files no worse than clones kept together [R12]. Göde and Koschke found that most clones are rarely changed at all [GK11]. Sjøberg and colleagues hired developers and logged their time across four systems and twelve named smells. Once file size and how often a file changed were held fixed, no smell cost measurable effort, duplication included [S13].

For twenty years these results were cited against each other. They are not in conflict. They measured different things, and the difference is the whole point. Kapser and Godfrey judged clones as they sat. Rahman counted defects per cloned line. Göde and Koschke counted how often clones were touched, and found: rarely. Sjøberg measured effort per *file*, and a file that contains a copy is not more expensive to edit than a file that does not. None of these studies asked what happens at the moment the duplicated *decision* changes, and Juergens's study, the one that did, found the cost there and only there. Eick's span, Mockus's diffusion, Eaddy's scattering and Hassan's entropy are all measures of the same moment, the change that has to reach several places, seen from the change's side rather than the clone's.

So the clone wars end in a sentence. Duplication costs when a duplicated decision changes, in the change that must find every copy and the fault when it does not, and it costs nothing you can measure while it sits. Count clones and you find nothing. Count effort per file and you find nothing. Count inconsistent changes, or faults at a sibling site, and you find it.

That sentence will matter in section 6, because the first attempt to measure Harbour counted clones and effort per file, found nothing, and said so.

## 4 The deletion that never comes

If the cost is paid when the decision changes, then the remedy is to leave one place for it to change in. Every migration pattern in the practitioner literature says this, and every one of them ends the same way. Hammant's branch by abstraction ends at step six: "Delete the first implementation" [H07]. Sato's parallel change ends in a contract phase, with a warning that "if the contract phase is not executed you might end up in a worse state than you started" [SA14]. Fowler's strangler fig ends when the host tree dies [F24]. The pattern is always: build the new path, move the callers, delete the old path. And the step that gets skipped is always the last one.

The record of how often it gets skipped is unusually good, because feature toggles and deprecated APIs leave a trail. In Chrome, a cleanup campaign triaged the release toggles and found that only one in five had actually been removed; of eleven marked "Removed", two were still in the code [RQ16]. Uber built a tool that generates the deletion diff and a bot that chases its owner, and even then most of the generated diffs landed only within days of a reminder [PIR20]. In a Smalltalk ecosystem a median of one project in five reacted to a deprecation at all, and parts of the ecosystem "remain in an inconsistent state for long periods of time" [RL12]. Nearly every Java client in a study of twenty-five thousand kept adding calls to APIs that had already been deprecated [SRB18]. Google's engineering book draws the practitioner's conclusion: "code is a liability, not an asset," and a deprecation that actually finishes is one with a deadline and a team staffed to see it through, because the advisory kind lacks enforcement and turns into whack-a-mole [G20].

The technical-debt literature says the same thing from inside the organisation. Cunningham's founding statement of the metaphor was about this exactly: "entire engineering organizations can be brought to a stand-still under the debt load of an unconsolidated implementation" [C92]. Besker, Martini and Bosch had 43 developers at six companies report their time weekly for seven weeks and found 23% of it lost to debt, with developers reporting that they were "frequently forced to introduce new" debt because of the debt already there [B19], which is the Perry and Wolf loop measured in people. In their multiple-case study Martini and colleagues name the specific failure: "non-completed refactoring," in which "the new API is added but the previous one cannot be removed," and a duplication that "was not considered as such (but only temporary duplication)" until the substitution turned out to be impossible for want of priority [MBC15].

Keep that phrase, *only temporary duplication*. It is what every comment in Harbour that says "keep in sync" and names a ticket means, and the Martini study was written about human engineers at five large companies in 2015.

## 5 The additive mind

Why is the last step the one that gets skipped? Psychology has a candidate answer, and it predates the agents. Adams and colleagues ran eight experiments in which people improved objects, essays, itineraries and structures, and found that people "systematically overlook subtractive changes" [AD21]: asked to make a thing better, they add to it, and reach for removal only when prompted, and sometimes not then. A 2026 follow-up found the same bias, stronger, in GPT-4 and GPT-4o [U26].

Then it was measured in code. Ebrahimi and colleagues gave language models bug-fix tasks whose correct patch deleted something, and found that the models carried out at most 71.7% of the deletions the developer's fix had made, even on tasks every model solved; among the patches that *passed the tests*, 29% kept the faulty code in place behind a guard, a pattern they call "Guard-and-Go" [EB26]. The paper's title is the thesis: to add is machine, to delete is human. A guard around a bug is a second decision about the same thing, one that disagrees with the first and now has to be kept consistent with it, and a passing test suite does not know.

The repository studies say agents do this at scale. Huang and colleagues, comparing agent-written pull requests with human ones in the same repository, found functions written by agents nearly twice as redundant with the code already there, and concluded that "LLM agents frequently disregard code reuse opportunities" [HU26]. Kashif and colleagues generated ten projects feature by feature with an agentic IDE and a person in the loop, found them 91% functionally correct, and found duplication the most common of 1,305 design issues [K26]. And He and colleagues, in a difference-in-differences design over hundreds of repositories that adopted Cursor and matched controls, found "a statistically significant, large, but transient increase in project-level development velocity" and "a substantial and persistent increase in static analysis warnings and code complexity," with the accumulated complexity "major factors driving long-term velocity slowdown" [HE26]. That is the closest published test I know of for the hypothesis this essay is about: the same repositories, the same tool, and the state of the code in one period predicting the output of the next.

Put the five sections together and the prediction is specific. An agent-built codebase will hold more decisions in more places than a human-built one, because agents re-derive instead of reusing and add instead of removing. The cost will not show up as clone counts or as effort per file. It will show up when a scattered decision changes, as the change that reaches some of the places, the review that notices, and the fault at the place it did not reach. And the ordinary remedies, asking for reuse and filing the cleanup, will not work, because they did not work for people either.

## 6 The measurement that fooled itself

Harbour is that codebase. Its production code grew four and a half times between May and October, nearly all of it written by agents, one ticket at a time. The paper that measured it got the answer wrong first and caught it the same day, and the way it went wrong is the way the clone literature went wrong for twenty years.

Harbour's archive records a decline. Correct, complete changes a week fell from 95 in June to 47 after mid-July, and the correct share of merged tickets from 84% to 72% [H4]. A series of earlier papers had attributed all of it to the process around the change: supervisors woken to be told nothing had changed, sessions re-reading what the last one read, plan-review loops. They set aside the escaped bugs that a later review had found in older code, on the grounds that those were old faults newly noticed rather than new faults, and concluded that "the cost rose while quality held, and it rose around the change, not in it" [H4]. The state of the code was not their question, and nothing in this essay disturbs what they measured. But the coherence paper took their conclusion as its starting point, and so its first version said there was no decline for coherence to explain. [H1] [H2]

Then it measured exposure at the file level, as Sjøberg had, and found no penalty, as Sjøberg had. And it controlled for the size of the change, as is standard. But touching a scattered decision is part of what makes a change large, so holding size fixed removed the effect it was looking for. [H1]

The correction came from asking what the set-aside bugs actually were. The reliability paper had set aside at least 67 of the 103 August and September escapes. Of the rows from those months that the coherence paper's own readers mark as found by a later sweep, 72% are coherence faults: a sibling the fix never reached, a copy that had drifted. The adjustment that made quality hold had removed, by construction, the rows where the hypothesis lives. "Quality held" was true of the faults users met, which stayed about a quarter coherence-caused throughout, and false of the stock, which the sweeps were finding as fast as they looked. Both attributions were right about what they measured. Tokens do go to supervision. And nobody had asked the review gate what it was sending things back *for*. [H1] [H2]

The lesson is the one from section 3, learned twice. Choose the clone's unit and you find nothing; choose the file's unit and you find nothing; set aside the faults found by the people looking for the thing and you find nothing. The cost of a scattered decision is only visible from the change, and only when the change is the one that had to reach every place.

## 7 The bill, itemised

Measured from the change, with classes written down in advance and readers who had not seen the first attempt, the bill is in three ledgers. [H1]

The first is the escaped defects. Of 190 bugs that got past review into the main branch, 86, which is 45%, have an unreconciled decision as their cause: most often a sibling left unfixed after one copy was fixed, next a change that landed in fewer places than it needed, then two copies that had drifted apart until the difference was the bug. The sweeps found what they found: "seven provider 409 sites missed LIN-2266's session clear," "nine addItem sites, only four resolve dispatchDefaults," "inline error envelope missed LIN-2351's 36-site fix." Among the bugs found by use rather than by a later sweep the share is 27%, and it has been about that in every month since the fleet started. Among the bugs a later review or sweep turned up in older code, it is 70%.

The second is the review gate. Of 397 verdicts since June that sent a plan or a change back, 143, which is 36%, name a sibling not updated, an existing path not used, or a divergence the change would make. The share rose from 22% in July to 47% in October. Read the verdicts and they say what Eaddy's theory said: "~48 sibling call sites broken live," "/recommend-and-dispatch drops periodicalId; /dispatch has it," "four description write paths, one tested."

The third is the change itself. Measured at the site rather than the file, 15% of PRs since June, about one in seven, touch a place where a known multi-site decision lives. Those PRs are five times the size at the median and their tickets are sent back two and a half times as often, more often in every size band. They do not escape more, and they do not take more hours at a given size. The size *is* the cost, or most of it: a change to a scattered decision is large because it has to be.

And the one measure that looked for Juergens's inconsistent change among the thirty-three *named* twins found almost nothing: of the edits that touched some of a twin's sites and not others, three were followed by a sibling fix. The sibling faults in the first ledger sit on decisions no comment ever admitted. The copies that cost are the ones nobody wrote down.

## 8 Four lives of a decision

Numbers across a codebase say how much. Following one decision from birth says how. The paper did that for four, walking every commit on the main branch. [H3]

The tolerant timestamp parse was born on 29 May as a private function in a store, returning nothing rather than throwing. It was copied into a second store in June, six more times in July under three different names, five more in August, and seven times in five days in October. One of those copies produced a bug: a caller passed a value the parser turned into a nothing that was not quite nothing, so a branch that should never run always ran. It was fixed at the one caller. No helper was ever proposed. The parser goes by three names, none of them exported, so no reviewer ever saw it as one thing.

The option bag was born in the routes on 13 June and spread to fourteen of them. On 17 July, when the hosting moved and the footer went blank, the function the bag calls was moved into its own module, and every one of the nine route copies kept the literal call. The same day a ticket noted that five of the copies' doc comments still said the old host's name. The stale documentation had travelled with the copies; plan review sent the fix back for the siblings it had missed, and the fix that finally landed rewrote seven files.

The hand-rolled workspace lookup was born on 20 January, in a direct push on the very day the canonical helper landed, because the renderers held an array and the helper wanted a session. It reached nineteen sites. The Drift & Coherence review tracked it through five editions, from four sites to sixteen to seventeen, and in August minted a ticket to change the helper's shape so the callers could use it. The ticket never landed. In September a test was added that pins the count of hand-rolled sites at sixteen. In October two agents added an eighteenth and a nineteenth site by naming the variable something the pin did not match, one of them with a comment saying that was why.

The inline error envelope was born on 11 January and had a canonical helper a week later, with twenty inline sites already behind it. There were 306 by the end of May and 357 by the middle of June. Then a sweep ticket deleted 304 of them in a day. Since then, forty new inline sites have been written, nearly all in new files, most of which import the helper. The flight-companion route imports the error helpers and wrote eight raw envelopes beside them.

Four decisions, one life. A private copy at birth; a copy with each new module; a helper that the callers do not adopt, because its shape does not fit or because nobody's ticket includes adopting it; a review that counts the sites and files a ticket nobody works; a pin that records the debt and is stepped around. And one sweep, the only consolidation that finished, which finished because it deleted.

## 9 Why asking does not work

Harbour tried the things one tries. Each of them is a way of asking, and each has a literature that says asking does not hold.

It ran a review that measures exactly the scattering described here, with five editions between June and September and a trend ledger. The ledger says the scattering worsened under it. The three fix tickets it minted sat in the backlog with no comments, and the review's own text says "the promotion path is still not converting." Its chosen remedy for the error envelope, adopt the helper on next touch and do not sweep, produced the flight-companion route. This is Google's advisory deprecation, and Robbes's ecosystem that reacts to a deprecation one time in five. [H1] [G20] [RL12]

It added parity tests, thirty-three of them, each asserting that two copies still agree. A parity test is honest. It is also a way of keeping the copies: the clearest case added a helper and a test that reads four files' source and asserts their literals match the helper's, and the four copies stayed, and a fifth arrived a week later outside the test. Martini's "only temporary duplication" has a test suite now. [H1] [MBC15]

It extracted helpers. Of four consolidations whose old pattern could be counted, one finished in its own PR; one left the decision at the callers because the helper takes the deciding value as an argument; one still has two of four old writers inline a month later; and one shipped with no importers at all and had to be adopted by a second ticket. Non-completed refactoring, as predicted: the new path is added, the old one is not removed. [H1] [MBC15]

And twice, it deleted. The meta-prompt, 106 kilobytes that had to be changed alongside the stage rules in 84% of the changes to those rules, was deleted outright in October, and there is now one place to change a stage's rules by construction. The error-envelope sweep in June deleted 304 sites in a day. In both cases the consolidation finished because nothing was left to converge. [H1]

The pattern across the four remedies is not that agents ignore instructions. It is that an instruction scoped to one ticket cannot reach a decision that lives in nineteen files, and a review that files a ticket has handed the work to a queue where, on the archive's count, half of all filings are never worked [H1]. The copies in Harbour were made with the original in view: seventeen of the thirty-three sit on a rule or a boundary, seven of them naming the rule in the comment. More asking does not change that.

## 10 What holds

Two things do, and both are structural.

The first is to leave one place. A shared pure module that the browser can load without a build step, which browsers have done for years, removes the reason for every server-and-browser twin at once. A layering rule with a named place for shared predicates removes the reason for the next four. Neither asks an agent to do anything; each takes away the reason a copy was reasonable. [H1]

The second is to change what finished means. Today a ticket is done when its acceptance criteria pass. A ticket that touches a scattered decision could instead be done when that decision has one authority again and the losing sites are deleted in the same change, or the ticket stays open. That is Google's compulsory deprecation with a deadline, scaled down to one change, and Harbour has not tried it. The paper specifies a six-week two-arm trial of exactly that rule, with its measures named in advance, and it is the next thing to run. [H1] [G20]

## 11 What the evidence does not show

Three things, and they should be said plainly.

It does not show a trend in the cost of a same-sized change. The hours a change takes are recorded only from September in the data the paper could reach. The coherence share of faults found by use is flat since June; the share of all faults rises from August, but it rises as the reviews began looking for siblings and the sweeps began, and a rise in finding is consistent with a stock that was there to find. Whether a change to Harbour in October costs more than the same change in July, for this reason rather than for the reasons the steady-base papers gave, is unmeasured. The paper says what would measure it. [H1]

It does not show how much the size control still hides. A change that touches a scattered decision is large because it must reach several places, so holding size fixed removes part of the cost, and what remains is the send-back gap. But a decision that lives in many places also lives in the hottest code, and the design cannot separate the scattering from the heat. [H1]

And it rests on readers. The counts were coded blind, against written classes, by readers who had not seen the first version; the independent check re-read forty of the escape rows and agreed on thirty-five. The readers are models of the same family as the author, and a class called "sibling-unfixed" states the hypothesis. The evidence quote on every row is the control, and the check is one; neither is proof. [H1] [H2]

What the evidence does show is enough to act on. The oldest argument in the field, the measured results once their units are read correctly, the new studies of what agents add and fail to delete, and the itemised bill from one agent-built codebase all say the same thing in the same place. A decision held in more places than a change reaches is a debt whose interest is paid at the gate and in the sweep. The remedies that ask for coherence have not held, in Harbour or anywhere the literature looked. The remedies that leave one place and delete the others held the two times Harbour tried them.

## Annotated reading list

Markers in the text refer to the entries below. Each entry says what the source is, what it supports here, and what it does not. A source from another field says so; an analogy is a mechanism to consider, never an estimate. Fifteen of the outside sources were checked against their primary text for the companion paper and the check's ledger is cited in the header; the rest were read as the paper's source list records, at the full text or abstract it names.

### The Harbour papers

- **[H1]** `coherence-as-it-grows.md` (version 3, 2026-10-08). Supports every Harbour figure in sections 1, 6, 7, 9, 10 and 11: the clone series, the thirty-three twins and the second reading (seventeen on a rule or boundary, seven naming the rule), the 86 of 190 escapes and their classes (42 sibling-unfixed, 27 missed-site, 12 drifted), the 143 of 397 send-backs (46% at plan review), the decision-level exposure table (211 of 1,385 PRs; median 229 against 45 production lines; send-backs 0.56 against 0.23 per ticket), the 158 subset edits and three sibling fixes, the four remedies in finding 5, the options and the two-arm trial. Does not support a trend in per-change cost before September, or an apportioning of the June-to-July fall; its Limits say so.
- **[H2]** `coherence-as-it-grows-check.md` (version 1, of the paper's version 2). Supports section 6's account of what the paper's first version got wrong and why, the "at least 67 of 103" rows the reliability paper set aside, and the forty-row blind re-read (35 agree, κ 0.75; on the checker's reading the 45% would be 44%). Does not check version 3's own changes.
- **[H3]** `coherence-as-it-grows-sources/lineage.md` and `.json`. Supports section 8 entirely. Ticket attribution for merges that name no ticket is from inner commits and marked as an estimate there; "never landed" for the helper-shape ticket is inferred from the absence of any citing commit; the lineage's own record shows the flight-companion route existing from July, so "new file" in section 8 means new to the error-envelope pattern, not new to the tree.
- **[H4]** The steady-base papers, read at the lines the header cites: `measuring-throughput.md` (95 and 47 a week), `why-throughput-halved.md` (84% to 72%), `reliability-baseline.md` (the rows set aside), `steady-base-menu.md` (the quoted sentence). They support the decline and its attribution in section 6. Nothing in this essay disturbs their measurements of where tokens go; the state of the code was not their question.

### The argument from design

- **[P72]** Parnas, D. L. (1972). On the criteria to be used in decomposing systems into modules. *CACM* 15(12). https://doi.org/10.1145/361598.361623. Supports the criterion in section 2 and its reading as a prediction. Argument with one worked example; no data.
- **[P79]** Parnas, D. L. (1979). Designing software for ease of extension and contraction. *IEEE TSE* SE-5(2). https://doi.org/10.1109/TSE.1979.234169. Supports "excessive information distribution" and the three-languages example. Argument from experience.
- **[P94]** Parnas, D. L. (1994). Software aging. *ICSE '94*. https://doi.org/10.1109/ICSE.1994.296790. Supports "ignorant surgery" and the two quoted consequences. Invited essay.
- **[PW92]** Perry, D. E., & Wolf, A. L. (1992). Foundations for the study of software architecture. *ACM SIGSOFT SEN* 17(4). https://doi.org/10.1145/141874.141884. Supports the drift loop. Definitions; no data.
- **[B75]** Brooks, F. P. (1975/1995). *The Mythical Man-Month*, ch. 4. Supports conceptual integrity as quoted. Essay from OS/360.
- **[B87]** Brooks, F. P. (1987). No silver bullet. *IEEE Computer* 20(4). https://doi.org/10.1109/MC.1987.1663532. Supports "no two parts are alike." Essay.

### The measured cost

- **[E01]** Eick, S. G., et al. (2001). Does code decay? *IEEE TSE* 27(1). https://doi.org/10.1109/32.895984. Supports the doubling of multi-file changes (under 2% to over 5%) and span predicting effort; fifteen years, one switch.
- **[MW00]** Mockus, A., & Weiss, D. M. (2000). Predicting risk of software changes. *Bell Labs Technical Journal* 5(2). https://doi.org/10.1002/bltj.2229. Supports diffusion as "essential to predicting failures"; same system; an industrial journal.
- **[E08]** Eaddy, M., et al. (2008). Do crosscutting concerns cause defects? *IEEE TSE* 34(4). https://doi.org/10.1109/TSE.2008.36. Supports the scattering–defect correlation and the quoted theory; three systems; correlations, not experiments.
- **[H09]** Hassan, A. E. (2009). Predicting faults using the complexity of code changes. *ICSE '09*. https://doi.org/10.1109/ICSE.2009.5070510. Supports entropy over churn; six projects.
- **[J09]** Juergens, E., Deissenboeck, F., Hummel, B., & Wagner, S. (2009). Do code clones matter? *ICSE '09*. https://doi.org/10.1109/ICSE.2009.5070547. Supports "nearly every second" unintentional inconsistent change being a fault, developer-confirmed; five systems. The hinge of section 3. Does not say how often clones change.
- **[KG08]** Kapser, C. J., & Godfrey, M. W. (2008). "Cloning considered harmful" considered harmful. *EMSE* 13(6). https://doi.org/10.1007/s10664-008-9076-6. Supports "as many as 71%" judged positive for maintainability, the authors' judgement per pattern; two C systems.
- **[R12]** Rahman, F., Bird, C., & Devanbu, P. (2012). Clones: what is that smell? *EMSE* 17(4–5). https://doi.org/10.1007/s10664-011-9195-3. Supports clones no more defect-prone, and dispersed no worse; four C systems; counts defects per cloned line, not inconsistent changes.
- **[GK11]** Göde, N., & Koschke, R. (2011). Frequency and risks of changes to clones. *ICSE '11*. https://doi.org/10.1145/1985793.1985836. Supports "most clones are rarely changed"; read at the abstract.
- **[S13]** Sjøberg, D. I. K., et al. (2013). Quantifying the effect of code smells on maintenance effort. *IEEE TSE* 39(8). https://doi.org/10.1109/TSE.2012.89. Supports the null at the file unit with hired developers and logged time; six developers, four systems. Its unit is the reason it found nothing, which is this essay's reading, not the authors'.

### The deletion that never comes

- **[H07]** Hammant, P. (2007). Branch by abstraction. https://paulhammant.com/blog/branch_by_abstraction.html. **[SA14]** Sato, D. (2014). ParallelChange. https://martinfowler.com/bliki/ParallelChange.html. **[F24]** Fowler, M. (2024). Strangler fig. https://martinfowler.com/bliki/StranglerFigApplication.html. Practitioner essays; support the shape of the pattern and its last step, not any rate.
- **[RQ16]** Rahman, M. T., Querel, L.-P., Rigby, P. C., & Adams, B. (2016). Feature toggles: practitioner practices and a case study. *MSR 2016*. https://doi.org/10.1145/2901739.2901745. Supports the Chrome campaign as corrected at the primary: of the release toggles triaged, 160 still in use, 51 (20%) removed, 44 lingering, two of eleven "Removed" still present; 39 releases.
- **[PIR20]** Ramanathan, M. K., et al. (2020). Piranha: reducing feature flag debt at Uber. *ICSE-SEIP 2020*. https://doi.org/10.1145/3377813.3381350. Supports the tool, the bot, 1,381 stale flags (17% of all), and 86% of generated diffs landing only within five days of a reminder; an industry report.
- **[RL12]** Robbes, R., Lungu, M., & Röthlisberger, D. (2012). How do developers react to API deprecation? *FSE 2012*. https://doi.org/10.1145/2393596.2393662. Supports the median 20% reaction and the quoted inconsistency; 2,600 Smalltalk systems.
- **[SRB18]** Sawant, A. A., Robbes, R., & Bacchelli, A. (2018). On the reaction to deprecation of clients of 4+1 popular Java APIs and the JDK. *EMSE* 23(4). https://doi.org/10.1007/s10664-017-9554-9. Supports 95–100% continuing to add deprecated calls; 25,000 clients.
- **[G20]** Wright, H. (2020). Deprecation. In *Software Engineering at Google*, ch. 15. https://abseil.io/resources/swe-book/html/ch15.html. Supports "code is a liability, not an asset"; the paraphrase of advisory versus compulsory deprecation follows its text ("often lacks enforcement"; "actively staffed by a specialized team through completion"; "a game of whack-a-mole"). Practitioner experience, no measurement.
- **[C92]** Cunningham, W. (1992). The WyCash portfolio management system. *OOPSLA '92*. https://doi.org/10.1145/157709.157715. Supports the founding quotation; experience report.
- **[B19]** Besker, T., Martini, A., & Bosch, J. (2019). Software developer productivity loss due to technical debt. *JSS* 156. https://doi.org/10.1016/j.jss.2019.06.004. Supports 23% and "frequently forced to introduce new TD" (abstract), with the cause, "due to the already existing TD," in the body; self-reported time, 43 developers.
- **[MBC15]** Martini, A., Bosch, J., & Chaudron, M. (2015). Investigating architectural technical debt accumulation and refactoring over time. *IST* 67. https://doi.org/10.1016/j.infsof.2015.07.005. Supports "non-completed refactoring" and both quotations; seven sites at five companies; qualitative, and about people, not agents.

### The additive mind and the agent studies

- **[AD21]** Adams, G. S., Converse, B. A., Hales, A. H., & Klotz, L. E. (2021). People systematically overlook subtractive changes. *Nature* 592. https://doi.org/10.1038/s41586-021-03380-y. Supports the bias as stated; eight experiments, none on software; an analogy to code, not a measure of it.
- **[U26]** Uhler, L., et al. (2026). Influence of solution efficiency and valence of instruction on additive and subtractive solution strategies in humans, GPT-4, and GPT-4o. *Communications Psychology* 4, 41. https://doi.org/10.1038/s44271-026-00403-0. Supports the bias appearing in the models and more strongly than in people; two pre-registered experiments, not on code.
- **[EB26]** Ebrahimi, A. M., et al. (2026). To add is machine, to delete is human. arXiv:2607.28887. Supports the deletion recall ceiling of 71.7% on solved tasks, the 29% of passing patches that are Guard-and-Go, and the name; preprint; benchmark tasks, not repositories.
- **[HU26]** Huang, Y., et al. (2026). More code, less reuse. *MSR 2026*; arXiv:2601.21276. Supports "nearly 1.87×" redundancy and the reuse quotation; the redundancy figure is from one repository, 3,858 PRs.
- **[K26]** Kashif, M., et al. (2026). Beyond functional correctness. arXiv:2604.06373. Supports 91% correct and duplication the most common of 1,305 issues; ten generated projects; preprint.
- **[HE26]** He, H., Miller, C., Agarwal, A., Kästner, C., & Vasilescu, B. (2026). Speed at the cost of quality. *MSR 2026*; arXiv:2511.04427. Supports the three quoted phrases and the design (806 adopters, 1,380 controls); duplicate-line density rose 7% and was not significant, which this essay does not lean on.

## Next

- **Run the paper's two-arm trial.** Six weeks, tickets that touch a listed seam assigned by parity to today's definition of done or to one that requires one authority and the deletions at close-out; sites per decision before and after, and the cost of later changes to the same seam. It is the one remedy this essay argues for that Harbour has not tried, and the paper has its measures written.
- **Apportion the June-to-July fall.** Put the runner's per-ticket working hours on a branch and set hours per same-sized change by month against decision-level exposure. Until that is done, section 11's first limit stands and this essay's claim is about where the cost lands, not how much of the decline it explains.
- **Name the unnamed decisions.** The 86 coherence faults and 143 coherence send-backs each name one; most are not in any census. A list of them, with their signatures, would turn the four lives in section 8 into forty.
