What is the stage selector given, and what does it actually need to choose correctly?
The selector is given more than it needs and less than it should, and cutting is not the fix. A call carries a median 11.9k characters (p90 28k, most 48k). Fifty-seven percent is fixed text (stage descriptions, rules and the reply contract), 24% the description, 12% the newest three comments and 4% facts computed in code. Those facts decide the routes that went wrong on 8 October. The description's late sections were not unseen: the hypothesis was tested and is refuted. The misroutes come from four shapes in the facts and the stage descriptions, and a fifth effect, the model splitting at temperature 0 wherever two signals disagree. Cropping the description or limiting comments saves 7–10% of the prompt and fixes none of them. Four small fixes, each at its cause, do (82 of 90 decision points right, against 39 of 90 on main), and they pass the existing gate. They are the recommendation for the next ticket; this paper changes no code.
Findings
1. The selector's input is small, mostly fixed, and the parent's description is not in it.
Over the 92 eval fixtures the prompt is 11.9k characters at the median. Stage descriptions are
6.5k of every prompt. Rules are 1.9k. The description as shown is 0.65k at the median but up to
35k; comment bodies 0.95k median, 8.4k most; trail and node facts 0.54k median, 1.1k most. On
the 8 October points the prompts were 11.9k to 43.9k characters, about 3k to 11k tokens. Code
cuts two things: a ## Implementation Plan over 3,000 characters is condensed (each subsection's
lead up to 400 characters, the whole session-fit section, and every line about session fit,
plan-review due or a revision), and comments are the newest three, each capped at 2,000
characters, plus the latest ruling and the latest person's comment wherever they sit. Everything
else in the description is shown whole. A task's parent arrives as identifier, title and state
only; its description never reaches the selector. All four routed surfaces build the same
context.
2. "Late sections of long descriptions go unseen" is refuted. Deleting the whole description
changed no route on P1–P7, P10, P11 or P12 (runs/descNone.json). On LIN-3357 and LIN-3358 the
deciding lines (plan-review due: yes, the session-fit line) are inside the condensed plan, at
line 213 of a 368-line prompt, and the model quotes them when it does choose plan-review.
Cropping the description's middle to its first and last 3k made LIN-3356's close-out worse: 10 of
12 answers became retrospective-audit, against 14 of 14 close-out with the full description.
3. Four shapes in the facts and stage descriptions, not length, produced the misroutes. Each was replayed at the moment of decision on the router model (counts pooled over valid batches on main):
- S1, defer against plan-review timing (LIN-3357, LIN-3358). The plan created its children
before its own plan-review, and
defer(rule 4) is checked beforeplan-review(rule 8), so the parent defers: 3 of 90 right after revisions. The FC's "plan" on these tickets was not the parent's answer. The parent saiddefer; the descent into LIN-3358's child LIN-3364 then saidplan(4 of 6), the last hop of the descent. - S2, a leaf whose plan lives on its parent (LIN-3354, LIN-3355, and the LIN-3364 hop). The
leaf reads "Implementation plan: none" and routes to
plan: 35 of 60 right. The plan is in the parent's description, which the selector never sees. - S3, conditional Approve read as a fix round (LIN-3356). "ledger: 4 item(s), 4 not discharged", with the review's findings in the comments, reads as work to do: 5 of 30 right.
- S4, merged but not Done (LIN-3356 while the close-out runs). "Merged" reads as
retrospective-auditthough the task is In Progress: 2 of 18 right.
LIN-3340 did not reproduce (66 of 66 right). The first LIN-3357 point (P1) stays wrong in every variant because the FC's own ruling said the plan was accepted; its gold is arguable.
4. Wavering is the model splitting at temperature 0. Where two signals disagree, identical
prompts give split answers, and the split moves between batches (LIN-3354 gave implementation
5 of 18 in one batch and 10 of 12 in another). "Asked twice within seconds, two answers" is this.
Comparisons therefore have to run their arms concurrently, and one batch of K=3 says little about
a boundary fixture.
5. Little is safe to cut, and it saves little. On the 92 fixtures, K=2 per variant, run concurrently:
| variant | right | mean prompt characters |
|---|---|---|
| main | 159/184 (86.4%) | 15,608 |
| no description | 150/184 (81.5%) | 11,906 |
| description cropped to its first and last 3k | 161/184 (87.5%) | 14,007 |
| newest comment only | 163/184 (88.6%) | 14,531 |
The description decides terse tickets (LIN-1892, FIX-830-neg, SYN-10, 12, 14, 23 fail without it). The cropped and comment-limited variants are within noise on the corpus but save 10% and 7%, and cropping made a known case worse (finding 2). Removing the facts mostly fixes P2–P6 (15 of 20), fixes P8–P9 (8 of 8) and breaks P11–P12 (0 of 8): the facts decide, both ways. Hiding comment bodies breaks P8 and P9 (a spike's PR link has no explanation). So no cut is recommended. Nor is there "plenty of safe context": the part that decides is the 4% of facts, and the 57% fixed text is the stage vocabulary.
6. Four fixes, each at its cause, repair the points and hold the gate.
| fix | cause | points alone | main |
|---|---|---|---|
| b | defer's "Not when" names plan-review when the task's own plan is due one (S1) |
30/30 | 3/90 |
| c | a leaf with no plan reads its parent's plan as its plan; providers pass the parent's description (S2) | 24/24 | |
| d | under an approving review the ledger reads "left for close-out to discharge" (S3) | 12/12 | |
| g | retrospective-audit's "Not when" adds "merged but not Done (close-out)" (S4) |
18/18 |
Together on the 15 points: 82/90 against 39/90 on main, same batch. On the gate
(scripts/eval/jev-routing-eval.mjs arm 3, openai/gpt-5.6-sol, K=3, 92 fixtures, arms
concurrent), b+c+d+g scored 241/276 against main's 240/276, 82 fixtures right by majority against
80, and none that main gets right goes wrong. That meets LIN-3300's rule. Rejected variants:
plan facts on nodes (no effect); withholding retrospective-audit from the options, which fixed
S4 but cost HAR-697 (1/9); a "Done: no" fact, which did not fix S4. Routing cost is unchanged
(input tokens within 1%).
Method
The harness rebuilds the exact routing prompt for a ticket as it stood at a past moment:
comments and the task's dispatch runs cut at that time, and description, state and child states
from the task-history snapshot taken at the next dispatch. It is buildSelectorArgs line for
line plus ablation knobs; with no knobs all 19 rebuilt prompts are byte-identical to the
library's. There are 15 decision points (P1–P12b, timed from comment and dispatch timestamps)
and 4 descent hops; gold is what the FC then dispatched. Calls use the workspace's router model at
temperature 0 with the same request body as getRecommendation. Ablations ran on the points
(K=4–6) and on the corpus, graded by the eval's own gradeAnswer. Candidate fixes lived in a
scratch worktree behind env toggles (fixes-experiment.patch; with toggles off the prompt
snapshots pass unchanged). The gate ran twice, main and fix concurrently. The population of
selector builders is closed by grep -rn buildRouterPrompt lib routes (only openrouter.js and
the snapshot test).
Limits
- The gate cannot see S1, S2 or S4. None of the 92 fixtures has those shapes, so the gate shows only that the fixes cost nothing elsewhere. The 82/90 on the decision points is the evidence they fix the misroutes, and those 15 points were the ones the fixes were designed against.
- The FC's own transcript is unavailable (
GET /flight-companion/transcriptsreturns 0), so which hop and call order the FC saw is inferred from replay. - K is small and the model is noisy. Boundary fixtures (HAR-697, SYN-12, SYN-15, SYN-22) need K≥6; a main-against-main gate moved 3 calls. The two fixtures that flipped to right are within that scale.
- The corpus ablation is K=2 and within noise; it supports "not worth cutting", not "cutting is harmful".
- Replay is of the 8 October code. LIN-3364 changed the run outcomes afterwards but left
blockedreading as "waiting on a person"; later changes are not assumed. - The router model is one model. A different router would split differently.
- LIN-3340 did not reproduce, so its two-answer behaviour is unexplained by this paper.
Next
The recommendation becomes the next ticket: ship b, c, d and g, and in the same change add the
15 points as eval fixtures, pay the prompt-size freeze, and re-run the gate at K≥6 on the boundary
fixtures. The selector-change pins that ticket must update are the
tests/fixtures/stage-router-prompts/ snapshots, tests/unit/prompt-size-budget.test.js (its
source-byte freeze fails on the patch), a scripts/prompt-template-change-log.md row, and
docs/architecture/prompt-system.md. Fix c touches Linear, Jira and Local providers; GitHub and
GitHub Projects have no hierarchy, so it is inert there. This paper itself updates none of them.
One line goes into proposals.md: whether the fixed 57% of the prompt (stage descriptions, rules and reply contract) is
what makes the model split where signals disagree.