What is the stage selector given, and what does it actually need to choose correctly?

The selector is given more than it needs and less than it should, and cutting is not the fix. A call carries a median 11.9k characters (p90 28k, most 48k). Fifty-seven percent is fixed text (stage descriptions, rules and the reply contract), 24% the description, 12% the newest three comments and 4% facts computed in code. Those facts decide the routes that went wrong on 8 October. The description's late sections were not unseen: the hypothesis was tested and is refuted. The misroutes come from four shapes in the facts and the stage descriptions, and a fifth effect, the model splitting at temperature 0 wherever two signals disagree. Cropping the description or limiting comments saves 7–10% of the prompt and fixes none of them. Four small fixes, each at its cause, do (82 of 90 decision points right, against 39 of 90 on main), and they pass the existing gate. They are the recommendation for the next ticket; this paper changes no code.

Findings

1. The selector's input is small, mostly fixed, and the parent's description is not in it. Over the 92 eval fixtures the prompt is 11.9k characters at the median. Stage descriptions are 6.5k of every prompt. Rules are 1.9k. The description as shown is 0.65k at the median but up to 35k; comment bodies 0.95k median, 8.4k most; trail and node facts 0.54k median, 1.1k most. On the 8 October points the prompts were 11.9k to 43.9k characters, about 3k to 11k tokens. Code cuts two things: a ## Implementation Plan over 3,000 characters is condensed (each subsection's lead up to 400 characters, the whole session-fit section, and every line about session fit, plan-review due or a revision), and comments are the newest three, each capped at 2,000 characters, plus the latest ruling and the latest person's comment wherever they sit. Everything else in the description is shown whole. A task's parent arrives as identifier, title and state only; its description never reaches the selector. All four routed surfaces build the same context.

2. "Late sections of long descriptions go unseen" is refuted. Deleting the whole description changed no route on P1–P7, P10, P11 or P12 (runs/descNone.json). On LIN-3357 and LIN-3358 the deciding lines (plan-review due: yes, the session-fit line) are inside the condensed plan, at line 213 of a 368-line prompt, and the model quotes them when it does choose plan-review. Cropping the description's middle to its first and last 3k made LIN-3356's close-out worse: 10 of 12 answers became retrospective-audit, against 14 of 14 close-out with the full description.

3. Four shapes in the facts and stage descriptions, not length, produced the misroutes. Each was replayed at the moment of decision on the router model (counts pooled over valid batches on main):

  • S1, defer against plan-review timing (LIN-3357, LIN-3358). The plan created its children before its own plan-review, and defer (rule 4) is checked before plan-review (rule 8), so the parent defers: 3 of 90 right after revisions. The FC's "plan" on these tickets was not the parent's answer. The parent said defer; the descent into LIN-3358's child LIN-3364 then said plan (4 of 6), the last hop of the descent.
  • S2, a leaf whose plan lives on its parent (LIN-3354, LIN-3355, and the LIN-3364 hop). The leaf reads "Implementation plan: none" and routes to plan: 35 of 60 right. The plan is in the parent's description, which the selector never sees.
  • S3, conditional Approve read as a fix round (LIN-3356). "ledger: 4 item(s), 4 not discharged", with the review's findings in the comments, reads as work to do: 5 of 30 right.
  • S4, merged but not Done (LIN-3356 while the close-out runs). "Merged" reads as retrospective-audit though the task is In Progress: 2 of 18 right.

LIN-3340 did not reproduce (66 of 66 right). The first LIN-3357 point (P1) stays wrong in every variant because the FC's own ruling said the plan was accepted; its gold is arguable.

4. Wavering is the model splitting at temperature 0. Where two signals disagree, identical prompts give split answers, and the split moves between batches (LIN-3354 gave implementation 5 of 18 in one batch and 10 of 12 in another). "Asked twice within seconds, two answers" is this. Comparisons therefore have to run their arms concurrently, and one batch of K=3 says little about a boundary fixture.

5. Little is safe to cut, and it saves little. On the 92 fixtures, K=2 per variant, run concurrently:

variant right mean prompt characters
main 159/184 (86.4%) 15,608
no description 150/184 (81.5%) 11,906
description cropped to its first and last 3k 161/184 (87.5%) 14,007
newest comment only 163/184 (88.6%) 14,531

The description decides terse tickets (LIN-1892, FIX-830-neg, SYN-10, 12, 14, 23 fail without it). The cropped and comment-limited variants are within noise on the corpus but save 10% and 7%, and cropping made a known case worse (finding 2). Removing the facts mostly fixes P2–P6 (15 of 20), fixes P8–P9 (8 of 8) and breaks P11–P12 (0 of 8): the facts decide, both ways. Hiding comment bodies breaks P8 and P9 (a spike's PR link has no explanation). So no cut is recommended. Nor is there "plenty of safe context": the part that decides is the 4% of facts, and the 57% fixed text is the stage vocabulary.

6. Four fixes, each at its cause, repair the points and hold the gate.

fix cause points alone main
b defer's "Not when" names plan-review when the task's own plan is due one (S1) 30/30 3/90
c a leaf with no plan reads its parent's plan as its plan; providers pass the parent's description (S2) 24/24
d under an approving review the ledger reads "left for close-out to discharge" (S3) 12/12
g retrospective-audit's "Not when" adds "merged but not Done (close-out)" (S4) 18/18

Together on the 15 points: 82/90 against 39/90 on main, same batch. On the gate (scripts/eval/jev-routing-eval.mjs arm 3, openai/gpt-5.6-sol, K=3, 92 fixtures, arms concurrent), b+c+d+g scored 241/276 against main's 240/276, 82 fixtures right by majority against 80, and none that main gets right goes wrong. That meets LIN-3300's rule. Rejected variants: plan facts on nodes (no effect); withholding retrospective-audit from the options, which fixed S4 but cost HAR-697 (1/9); a "Done: no" fact, which did not fix S4. Routing cost is unchanged (input tokens within 1%).

Method

The harness rebuilds the exact routing prompt for a ticket as it stood at a past moment: comments and the task's dispatch runs cut at that time, and description, state and child states from the task-history snapshot taken at the next dispatch. It is buildSelectorArgs line for line plus ablation knobs; with no knobs all 19 rebuilt prompts are byte-identical to the library's. There are 15 decision points (P1–P12b, timed from comment and dispatch timestamps) and 4 descent hops; gold is what the FC then dispatched. Calls use the workspace's router model at temperature 0 with the same request body as getRecommendation. Ablations ran on the points (K=4–6) and on the corpus, graded by the eval's own gradeAnswer. Candidate fixes lived in a scratch worktree behind env toggles (fixes-experiment.patch; with toggles off the prompt snapshots pass unchanged). The gate ran twice, main and fix concurrently. The population of selector builders is closed by grep -rn buildRouterPrompt lib routes (only openrouter.js and the snapshot test).

Limits

  • The gate cannot see S1, S2 or S4. None of the 92 fixtures has those shapes, so the gate shows only that the fixes cost nothing elsewhere. The 82/90 on the decision points is the evidence they fix the misroutes, and those 15 points were the ones the fixes were designed against.
  • The FC's own transcript is unavailable (GET /flight-companion/transcripts returns 0), so which hop and call order the FC saw is inferred from replay.
  • K is small and the model is noisy. Boundary fixtures (HAR-697, SYN-12, SYN-15, SYN-22) need K≥6; a main-against-main gate moved 3 calls. The two fixtures that flipped to right are within that scale.
  • The corpus ablation is K=2 and within noise; it supports "not worth cutting", not "cutting is harmful".
  • Replay is of the 8 October code. LIN-3364 changed the run outcomes afterwards but left blocked reading as "waiting on a person"; later changes are not assumed.
  • The router model is one model. A different router would split differently.
  • LIN-3340 did not reproduce, so its two-answer behaviour is unexplained by this paper.

Next

The recommendation becomes the next ticket: ship b, c, d and g, and in the same change add the 15 points as eval fixtures, pay the prompt-size freeze, and re-run the gate at K≥6 on the boundary fixtures. The selector-change pins that ticket must update are the tests/fixtures/stage-router-prompts/ snapshots, tests/unit/prompt-size-budget.test.js (its source-byte freeze fails on the patch), a scripts/prompt-template-change-log.md row, and docs/architecture/prompt-system.md. Fix c touches Linear, Jira and Local providers; GitHub and GitHub Projects have no hierarchy, so it is inert there. This paper itself updates none of them. One line goes into proposals.md: whether the fixed 57% of the prompt (stage descriptions, rules and reply contract) is what makes the model split where signals disagree.