1. The spec is the equaliser
With a spec, the smallest and largest models landed within two runs of each other. Without one, the spread ran from 7 of 15 to 0 of 15. If you can only change one thing about a risky ticket, the spec moves the result more than the model choice does.
2. Benchmark order is not safety order
Anthropic reports Sonnet 5.5 at 70.6% on Terminal-Bench 4.0, above Opus 5.5 at 66.4%. On the migration with no written rules, Sonnet 5.5 dropped users.name in all three runs; Opus 5.5 did it once. Benchmarks reward finishing the task; this failure is finishing it too thoroughly.
3. Small models skip the README
With the README stating that the billing sync reads users.name, Opus 5.5 and Sonnet 5.5 kept the column in 6 of 6 runs. Haiku 4.5 dropped it in 3 of 3. On the coupon task it failed both coupon-reuse checks in 3 of 3 runs. Rules in the repo protect you only if the model goes and reads them; rules in the prompt reached all four.
4. A cheap model with a spec is not cheap
Haiku 4.5 with a spec averaged 15 to 40 turns, 129 seconds and $0.20 a run, about the same cost as Opus 5.5 with a spec ($0.22) and twice as slow. Sonnet 5.5 was the cheapest way to a passing run: $0.11 and 33 seconds on average with a spec.