Can a cheap model implement Harbour tickets to Opus's review standard, and what does it cost?

Yes, on small tickets, and the review is the cost. Thirteen tickets were implemented by three OpenRouter models through opencode over an evening and a morning, 12 to 13 September 2026, each judged by an Opus review. Eight passed first time and every one passed within one round. No reviewer found a logic defect; the five send-backs were all second-order, a layout, a docblock claim, a missing test, stale prose, a monitor that reports clean on an empty fetch. Each Opus review cost $2.26 to $4.79 API-equivalent regardless of the ticket's size, so the 22 reviews of the round cost about $62, which is about one and a half points of the weekly subscription window; the thirteen Opus close-outs that then merged every PR cost another $52. The implementations cost about $3.75 a ticket on GLM and $0.45 on Flash, on the operator's OpenRouter export, a reading Harbour's own relay cannot make, so the subscription's share of a cheap-implemented ticket is two Opus sessions, about $7.50, and the implementer's share is noise. A floor run afterwards found the bottom of this harness: the older DeepSeek V4 Flash produced two green PRs for a tenth of a cent each, gpt-oss-120b reported done or failed inside a minute without touching the tree, and an 8B Llama hung until aborted. The seven launch failures in the round were a daily spend cap, a stale model catalog, and CPU starvation, none of them the model. An evening of one-ticket experiments then pushed the boundary outward: Flash implemented an hour-plus provider-and-docs ticket, a thin ticket with no plan, and a client-side page ticket both single-shot and stepped, all four approved by Opus within one round and all four merged and set Done by Flash close-outs after a human read each ledger, so the subscription's share of those tickets was one Opus review each. Flash research on two tickets went unfaulted by Opus plan-review; Flash plans on the same two went round twice without clearing, on the class bound each time. A run the next morning, driven by an Opus autopilot session rather than conned by hand, gave Flash six harness-code tickets: one merged, two research passes cleared the gate, and two implementation tickets went round three times each on green-CI legs that were wrong against the real rows. The driver cost what conning had, about $32, and research would have saved every extra round.

Findings

Every ticket produced a reviewable PR on its first session, and none looped. Thirteen tickets, thirteen PRs with green CI, from three models the runner had never run before, with no human touch between dispatch and verdict. Harbour's own five-hour session watchdog and error-recovery pathology, the failure mode the dash-analysis found in three quarters of sessions, did not appear once.

First-pass rates tie between GLM and Flash; both reach 100% after one round.

Model Tickets First-pass Within one round Median implementation
z-ai/glm-5.3 6 4 6 20 min
deepseek/deepseek-v4.1-flash 5 3 5 14 min
google/gemini-3.8-flash 2 1 2 24 min

The routing proposal's Sonnet baseline is 84% first-pass over 115 cases. Thirteen is too few to place either model against it, but eight of thirteen first-pass with every miss recovered in one round is not obviously worse.

The misses are all second-order, and they are the misses a test would not catch. LIN-2830: a grid with 792 px of fixed columns inside a 263 px card, clipped by the site's overflow rule, in a change whose logic and tests were correct. LIN-2647: a header rewritten to remove an overclaim that introduced a new overclaim in the same sentence. LIN-2575: a correct change at four sites, one of which, the live path, no test pinned. LIN-2838: a rule changed in code and left standing in four places of prose, one served to agents. LIN-2573: a new scheduled scan that reports clean when it fetches nothing. Opus found each with a measurement, not a reading: rendered widths, a five-minute clock advance, a full-suite mutation, a prose census, an empty-fetch run.

Review cost is flat with size, so the saving scales with the ticket. Harbour prices Opus reviews from their relayed tokens at $2.26 to $4.79, five to nine minutes each. The one-line CLAUDE.md change cost $2.54 to review; the 414-line readout change cost $4.79. The review regrounds, mutation-checks and writes a ledger whatever the diff, so on a ten-minute ticket the review is most of the cost, and on a two-hour ticket it would be a fraction. A bake-off of small tickets, chosen small for safety, understates what a cheap implementer saves.

Harbour cannot see what an opencode session cost. The runner posts the usage of the one message it awaited, which is the final step (opencode-runner.js:383-387), so the relayed figure for a thirty-minute GLM session was $0.027 against $3.45 on the OpenRouter meter, and Flash sessions relayed under a cent. Filed as LIN-2835. The operator's hourly export for the window (12 September 19:00Z to 13 September 09:59Z) reads GLM $34.68 across nine implementation sessions, one blocked run and thirteen shadows, Flash $2.27 across eight sessions, Gemini $7.07 across four. Split by session-minutes, GLM is about $3.75 a ticket and $0.95 a shadow, Flash $0.45 a ticket, Gemini $3.55. Flash matched GLM's approval rate at an eighth of its price.

A cheap reviewer agrees with Opus when the PR is clean and misses the second-order finding cold. Thirteen shadow reviews on GLM used the identical prompt Opus received, with a header naming them shadows and forbidding any write beyond one comment. Eleven agreed in full. One (LIN-2760) found the same case Opus raised as a ruling for the operator and blocked on it instead. One (LIN-2573) approved, before Opus posted, the scan that reports clean on an empty fetch. The two Request Changes agreements (LIN-2575, LIN-2838) were formed with the Opus comment already on the thread, because the proxy's issue read returns every comment; the shadow said so itself. The floor run added two more cold pairs: on LIN-2697 both approved, and on LIN-2637 the shadow asked for the same missing test Opus asked for, a second-order finding, formed before Opus posted. So the cold evidence is now two agreements and one miss, the miss on the hardest finding of the round, and the warm evidence is that GLM reproduces Opus's findings when it can see them.

Close-out costs what review costs, so the subscription pays twice per ticket. The thirteen approved PRs were merged by thirteen Opus close-outs on 13 September, dispatched in batches of four with nothing else on the queue. Every one merged and set Done in a single session; none was sent back, none hit a merge conflict, and main stayed green through the sequence on both repositories. They cost $1.62 to $8.88 each, $51.95 in all, with a median of nine minutes; the three over $5 (LIN-2575, LIN-2645, LIN-2573) were the ones with non-empty ledgers to discharge on the landed commit. With the reviews, that is about $114 of Opus per thirteen tickets, or $8.80 a ticket, against an implementer cost of $0.45 to $3.75. The cheap implementer removes the one session whose price scaled with the work and leaves the two whose price does not.

One cheap close-out on an empty ledger merged correctly; one Opus close-out finished and then never said so. After the floor run, LIN-2697's close-out (ledger explicitly empty, both reviews Approve) was given to DeepSeek V4 Flash 0731 on opencode, with LIN-2637's close-out on Opus as the control. Flash merged PR #1485, verified the landed commit, re-ran the pinned suite, set Done, posted a summary of the house shape and filed one follow-up for a class sibling the review had named as outside scope, in 13 minutes for a tenth of a cent. Opus merged PR #1486 and set Done within three minutes for $3.44; the dispatcher recorded that session COMPLETED, but Harbour never received its terminal marker, so the dispatch read as still running for over three hours until it was aborted by hand. One ticket each; the Flash result is a reason to run the empty-ledger case again, not to change the default, and the lost marker is a relay fault of the LIN-2835 kind, not a model one.

The floor of this harness is the older DeepSeek Flash; two cheaper models did no work. Five more tickets of the same shape ran after the close-outs: two each on deepseek/deepseek-v4-flash-0731 and openai/gpt-oss-120b, one on meta-llama/llama-3.1-8b-instruct as a control expected to fail. DeepSeek V4 Flash produced two PRs with green CI (LIN-2697 in 10 minutes, LIN-2637 in 28), one approved first time with an empty ledger and one sent back for a single missing test, fixed in a test-only commit and approved on re-review, so two of two within one round; relayed usage for all three sessions was under a third of a cent. gpt-oss-120b posted [done] on LIN-2504 after six seconds and fifteen output tokens with no tool call, and [failed] on LIN-2708 after a minute, claiming that public/observation.js and the function the ticket names do not exist; both exist on main. Llama 3.1 8B sat at a frozen ten-minute heartbeat for 49 minutes and was aborted, with no usage relayed at all. The two zero-work results cost under a tenth of a cent each, which is the point: below some capability line the failure is not a bad PR but a confident non-event, and Harbour's terminal markers cannot tell it from a finished ticket without reading the feedback.

The failures were the plumbing's, and every one was invisible from Harbour. Seven launches died: three to a $20 daily spend cap on the OpenRouter key (12 September from 20:50Z, found the next morning), two to opencode's cached model catalog refusing the three-day-old deepseek-v4.1-flash id while a server with a fresh cache accepted it, one to a 20-second readiness bound missed while a claude-code cold start pushed host load from 2 to 16 on ten cores, and two shadow sessions that hung after the cap and were aborted at a hundred minutes. LIN-2839's research, an Opus session reading the host's own logs, found no concurrency ceiling: twelve opencode servers ran side by side with 120 ms boots, and the conning session's working cap of four was a selection artefact. Harbour saw an exit code for each; the actionable text sat in a per-session log nobody reads (LIN-2837).

A model that reads "no code changes" and parks is worth recording. LIN-2415, dispatched in error, forbids code. GLM confirmed the prerequisite deploy, found the real cause of the missing stamp, a trailing newline, wrote a runbook and parked BLOCKED in four minutes. The fix it implied became LIN-2838, implemented by Flash and approved the same morning.

The boundary is not where the bake-off left it. Four one-ticket experiments the same evening, each judged by Opus review with a cold GLM shadow. LIN-2787, an hour-plus ticket across the Linear provider, a proxy route, the docs contract and a regression test, on Flash 0731 after two launch-time HTTP 500s on Flash V4.1: Approve, conditional, both reviewers finding the same inherited pagination item cold. LIN-2121, a thin ticket with no plan, on Flash V4.1: Approve, conditional, one hard gate item on a data path that the same session fixed with the reviewer's own probe as its test. LIN-2772, a client-side page behaviour ticket single-shot on Flash 0731: Approve, conditional, with the shadow stricter than Opus for the first time, on an unpinned rollback exit the session then pinned. LIN-2771, its sibling from the same review, stepped in four beats with the conning session at the wheel: Approve, conditional on the same two surfaces both reviewers named. Every send-back-shaped finding was closed in-session by a follow-up for pennies. Nothing here says where the boundary is; it says it is further out than one ticket per shape can see.

Cheap close-out held on non-empty ledgers when a human read them first. Four more Flash close-outs after LIN-2697: LIN-2772, LIN-2121, LIN-2787 and LIN-2771, each on a ledger with inside items, each after the conning session read the ledger, accepted the items on the ticket with the precondition each named, and had the in-session follow-up close the rest. Each merged, verified on the landed commit, set Done, cited the acceptance item by item and filed only outside items; LIN-2787's ran the review's post-deploy check on the live site. Two to ten minutes and under a cent each, against $1.62 to $8.88 on Opus. The condition is the whole finding: the cheap close-out did not judge, it cited.

Cheap research passed the gate; cheap plans did not, twice. Flash 0731 wrote research on LIN-2153 and LIN-2322 in under 25 minutes each for a third of a cent, with file-and-line citations across both repositories, the papers read, and on LIN-2153 the finding that the proposal collides with the manual's own doctrine. Neither Opus plan-review faulted it. The Flash plans built on it were sent back twice each, both times on the class bound rather than a member, with six of seven checks passing on the second round. The gate sends 88% of plans back once and a third back twice, so two rounds is inside the population; the difference from Sonnet, if there is one, is not visible at two tickets. One second-round plan session hung for an hour and was aborted; its retry finished.

The harness failed where the models did not. With seven or eight opencode servers on one host, three launches died before any tool ran, two with an HTTP 500 on the first message and one with the server not ready in twenty seconds, and one session hung. Every retry worked. The runner's five-minute stall note carries a fixed placeholder session id in every instance, so it is a template, not a signal; the duplicate-dispatch guard refuses a retry for five minutes after a launch failure and force does not override it on the recommend-and-dispatch verb.

Cheap implementation of harness code was green and wrong, four times in six. The 14 September run (LIN-2875) gave Flash 0731 three tickets whose subject is the harness's own feedback: a stalled note's id (LIN-2873), a guard over prior dispatches (LIN-2872 with LIN-2869), and a state write on a failed turn (LIN-2837). Six implementation legs; five green on CI with mutation witnesses; four wrong. LIN-2872's first leg returned false on all three incident rows the ticket was filed for, its second still did, and its third, given the criterion as an allow-list rather than a symptom, fixed the three rows and exempted every failed opencode row instead. LIN-2837's first leg reverted a guard from LIN-2556, its second guarded one arm and left the sibling, and its third fixed exactly what it was told to, after which the review found the fourth member of the class in a test. Every defect was found the same way and never by CI: the Opus reviewer, and then the driver, fetched the shipped predicate and ran it against the real stored feedback, or diffed the branch's behaviour against base. The tests encoded the same assumption as the code about which harness lines mean work, so a mutation witness proved the tests pinned the code and nothing about whether the code matched the harness. On the one leg that had an explicit criterion Flash was right first time about the thing the criterion named and wrong about the thing it did not; the driver's reading, which this paper adopts, is that the failure mode is under-specification, not capability. That distinction is exactly what the next finding is for.

Research first, always, on cheap workers. Three tickets in that run were judged by the conning session to be understood to the line, and research was skipped on all three. Each then cost at least one review round on something research asks for by name: LIN-2873's ticket asserted a root cause the code contradicts, which the staleness check exists to catch; LIN-2872 needed the class of markers a failed opencode row carries enumerated before anyone wrote a predicate over them; LIN-2837's two-arm guard is the sibling sweep. Flash research on the same harness cleared the Opus gate three of three times on 13 and 14 September (LIN-2153, LIN-2322, LIN-2835), in 12 to 22 minutes for under a cent each, and LIN-2835's found four live opencode servers and proved its answer against them. A failed review round costs an Opus review at $3 to $4, two or three driver wakes at about $3 each, and forty to fifty minutes. Research pays for itself at one round saved in twenty; on this run it would have been three of three. Flash plans stay out of the loop on the evidence so far; Flash research goes in front of every cheap implementation, and the exemption for a ticket that looks understood is withdrawn.

A driver at the top costs what conning costs, for now. The run's orchestrator was an Opus autopilot session launched from a brief on LIN-2875 and observed by the conning session, which read ledgers and answered gates on the tickets. The driver's discipline was the best seen on Harbour: it verified every worker claim against the artefact rather than the report, waited at every human gate, declined an instruction it had evidence against, and stopped lanes at the brief's conditions. It also woke about seventeen times, and each wake re-read its whole history: fifty million cache-read tokens against 150 thousand output, about $32 at Harbour's Opus 5 table (lib/model-pricing.js), against about $35 for the hand-conned evening before. The whole run was about $57 API-equivalent on the subscription and two cents relayed from OpenRouter: five Opus reviews at $3.70 each, one Opus research at $6, the driver, and Flash. The saving the driver buys is the conning session's attention, not tokens, until its context is compacted between wakes (LIN-2117) or the cloud Flight Companion takes the seat.

The silent hangs were a permission prompt. Four opencode sessions across the two days heartbeat as busy for ten to eighty minutes with no output. The driver read the per-session log of the first one and found it ends on permission=external_directory: the worker had read outside its clone, opencode asked, and in a headless run nobody answers. A sweep of all twelve sessions on the host found exactly one unanswered prompt, the wedged one. A close-out then wedged the same way while grepping the host's state directory for a proxy token it already held, so the trap is not about the task's subject. One line at the top of the prompt, stay inside the clone, prevented it in both later workers. The Opus research on LIN-2874 also settled the other silent case: the LIN-2637 close-out did post its terminal marker, and Harbour returned 502 to the entire finalize tail inside 174 milliseconds, the only seven failed feedback posts in two days of oplog. Neither was the model. Both are on the board (LIN-2876, LIN-2879), with the reaper that deleted the one log that could have answered the 13 September hang (LIN-2877).

Method

Twelve small tickets whose descriptions were already the plan, plus one filed during the run, were ratified on LIN-2828: six for GLM, four for Flash, two for Gemini as a control, one bonus for Flash. No Todo ticket in the workspace carried a reviewed plan, so each was dispatched straight to implementation with the verb pinned, from a conning session over the read-write proxy: POST /api/proxy/recommend-and-dispatch with kind: implementation, harness: opencode, the model id, target: cli, appendProxyContext: true and sessionId: LIN-2828-bakeoff. On [done] an Opus review was dispatched the same way with kind: review and no model, so the workspace defaults (claude-code, opus, medium) applied. The review's stored prompt was read back from GET /dispatch/{id}/prompt, its access block cut, a shadow header prepended, and the result dispatched as kind: custom on GLM. A Request Changes verdict was answered by a fix round on the same model and a re-review. Verdicts are the DONE: line of each review's feedback and the review comment on the ticket. Durations and review prices are Harbour's /cost lineages, read 09:25Z on 13 September. The results table is LIN-2831's description; every event is a comment on LIN-2828.

Close-outs were dispatched the same way with kind: close-out and no model, after the operator's pause and decision on 13 September: LIN-2838 first so LIN-2415 could be verified on the deploy, LIN-1220 alone because a merge to simple-dispatcher's main restarts the dispatcher through its watcher, then the remaining eleven in three batches. The floor run used the same implementation dispatch with the three floor model ids, an Opus review only where a PR reached green CI, and the shadow dispatched within seconds of the Opus review so its verdict formed before any Opus comment existed. The Llama session was aborted over the proxy after 49 minutes at an unchanging heartbeat.

Limits

Thirteen results, all small, none over an hour, chosen by the person running the bake-off. Per-ticket implementation cost is an hourly total split by session-minutes, because the relay is wrong by two orders; GLM's implementation and shadow sessions share one bucket. The shadow reviewer is the same model as one of the implementers, and eleven of its thirteen verdicts were on PRs Opus also approved, where agreement is cheap; its one cold test was a miss. Gemini had two tickets. The upstream host and quantisation behind each session was not recorded. The failure-rate figure for the harness is inflated by one configuration fault (the cap) and one catalog race that a week-old model id would not have hit. The floor run is two tickets per model, and it says nothing about why gpt-oss-120b did not act or why the 8B model hung: opencode relays neither the tool trace nor the provider's response for either, so the cause could be the model, the harness's tool-calling on that model, or the upstream host, and this run cannot tell them apart. Close-out costs are Harbour's own relayed Opus tokens and are comparable with the review costs but not with the OpenRouter export.

Next

  • Research first on every cheap-worker ticket for the fleet week, and count review rounds against this run's three-per-ticket on the two harness tickets.
  • Land LIN-2876 and LIN-2835, then re-run the LIN-2874 research on Flash with the same prompt, to learn whether the permission fix alone makes the host readable.
  • Price the driver again with beat-boundary compaction (LIN-2117), and again from the cloud Flight Companion when it can dispatch; the number to beat is $32 for six tickets.
  • Give a harness-code ticket to Sonnet, to learn whether green-and-wrong is the model or the shape.
  • Run the stepper pair the other way round, the easier ticket stepped, and on a shape that failed single-shot, before reading anything about stepping from one pair.
  • Give Flash a plan on a ticket it has just researched, with nothing landing under it.
  • Finish LIN-2872 from an Opus plan over the review's six ledger items, and decide whether its correct half (LIN-2869) lands first.