Same tickets, four Claude models: what a spec is worth at each size

We re-ran the three agent-run cases on Sonnet 5.5, Haiku 4.5 and Fable 5.1, with the same fixtures, prompts and frozen hidden suites as the original Opus 5.5 runs. The spec made the models nearly interchangeable. Without it, the differences were large, and they did not follow the benchmark rankings.

With a spec28 of 30 runs passed every hidden check
Without a spec12 of 48 runs passed every hidden check
Dropped a column billing readsSonnet 5.5: 3 of 3 runs with no written rules

Runs that passed every hidden check

Each cell is runs with a perfect hidden-suite score out of runs recorded. "No docs" and "No rules" are the one-line prompt in a repo with its written constraints removed; "Docs" and "Rules" are the one-line prompt with them present. Fable 5.1 got one run per variant and no no-docs runs, so treat its column as a spot check.

Case and variantOpus 5.5Sonnet 5.5Haiku 4.5Fable 5.1
Coupon, vague2/33/30/31/1
Coupon, spec3/33/33/31/1
API errors, no docs0/30/30/3not run
API errors, docs0/30/30/30/1
API errors, spec3/33/33/31/1
Migration, no rules2/30/30/3not run
Migration, rules3/31/30/30/1
Migration, spec3/33/31/31/1
All spec runs9/99/97/93/3
All non-spec runs7/154/150/151/3

"Perfect" is strict: the API "docs" runs mostly failed a single check, one error shape across all endpoints. The per-check detail for every run is in the run packs.

Four findings

1. The spec is the equaliser

With a spec, the smallest and largest models landed within two runs of each other. Without one, the spread ran from 7 of 15 to 0 of 15. If you can only change one thing about a risky ticket, the spec moves the result more than the model choice does.

2. Benchmark order is not safety order

Anthropic reports Sonnet 5.5 at 70.6% on Terminal-Bench 4.0, above Opus 5.5 at 66.4%. On the migration with no written rules, Sonnet 5.5 dropped users.name in all three runs; Opus 5.5 did it once. Benchmarks reward finishing the task; this failure is finishing it too thoroughly.

3. Small models skip the README

With the README stating that the billing sync reads users.name, Opus 5.5 and Sonnet 5.5 kept the column in 6 of 6 runs. Haiku 4.5 dropped it in 3 of 3. On the coupon task it failed both coupon-reuse checks in 3 of 3 runs. Rules in the repo protect you only if the model goes and reads them; rules in the prompt reached all four.

4. A cheap model with a spec is not cheap

Haiku 4.5 with a spec averaged 15 to 40 turns, 129 seconds and $0.20 a run, about the same cost as Opus 5.5 with a spec ($0.22) and twice as slow. Sonnet 5.5 was the cheapest way to a passing run: $0.11 and 33 seconds on average with a spec.

Cost and time per run

Averages as reported by Claude Code (API-equivalent cost). Turns are the range across runs.

ModelVague promptSpecTurns with a spec
Opus 5.5$0.18 · 47 s$0.22 · 63 s3–6
Sonnet 5.5$0.09 · 26 s$0.11 · 33 s3–6
Haiku 4.5$0.10 · 59 s$0.20 · 129 s15–40
Fable 5.1$0.71 · 90 s$0.83 · 95 s4–5

The two spec runs that failed

Haiku 4.5, migration S2: names reversed

The spec said the last token is last_name and everything before it is first_name. The build stored Mary Ann Smith as Ann Mary / Smith. Every other name was right, which is why a spot check would miss it and the "no characters lost" check caught it.

Haiku 4.5, migration S3: backfill outside the migration

Migration 003 only adds the two columns. The backfill runs in application code inside open(). A deploy that runs migrations separately, or any other service using the database, sees empty columns. The spec asked for the migration to backfill; the build's own tests passed because they open the database through open().

Fable 5.1, briefly

One run per variant, so no more than a spot check. With a spec, Fable 5.1 passed all three cases. Without one it matched Opus 5.5's pattern: coupons right, one error shape missing on the API, and on the migration it kept users.name but changed createUser and renameUser to take { first_name, last_name } instead of name, so every existing caller that passes name now creates users with no name at all. At $0.71 to $0.83 a run it was the most expensive model here by a wide margin.

Choosing a model for spec-first work

Spec first, then pick the model

Once the rules are in the prompt, Sonnet 5.5 passed everything here at half Opus's cost and time. Without the rules, no model was safe on the migration.

Put hard constraints in the prompt for small models

A README line was enough for Opus and Sonnet and not for Haiku. If you route routine tickets to a small model, paste the readers-and-boundaries section into the task itself.

Test at the boundary that matters

Both failed spec runs passed their own tests. A check that runs the migration the way production does caught them. See test evidence gates.

Method and limits

  • Claude Code 2.1.284, headless, one prompt per run, fresh fixture copy, no user settings, hooks or MCP servers. Models: claude-opus-5-5 (runs recorded September 28, 2026), claude-sonnet-5-5, claude-haiku-4-5-20251001 and claude-fable-5-1 (September 29).
  • Fixtures, prompts and hidden suites are identical to the original runs; nothing was tuned per model.
  • Three runs per variant (one for Fable 5.1) shows what can happen, not how often. Only Claude models were tested, through one agent; other vendors' agents read different instruction files (see what each agent reads).
  • Total cost of all 78 runs: about $14.84 as reported by Claude Code.

Editorial note

Every number on this page comes from the recorded runs in the run packs. We fixed nothing in the agents' output before scoring it. Benchmark figures are Anthropic's, as published on its model pages.