Do specs still matter with Opus 5.5? 24 recorded agent runs

We gave Claude Code three ordinary tickets, a coupon feature, an error-format cleanup and a schema change, and ran each one several times: from a one-line prompt, from a spec, and in some cases with the repo's written rules removed. Hidden test suites, frozen before the first run, scored every result. Every prompt, diff and score is downloadable.

Spec runs9 of 9 passed every hidden check
Non-spec runs7 of 15 passed every hidden check
Broke something outside the repo3 runs, all in repos with the written rules removed

Three findings

1. Vague prompts mostly work

On small, readable repos the model fixed crashes, closed stack leaks, capped discounts and backfilled data correctly from one sentence. If your review question is "does it run", a current model often passes it.

2. Written constraints stop breakages

A 13-line mobile contract doc and three README lines were enough to stop the breaking changes, even with a one-line prompt. Without them, two of three API runs broke the shipped app and one migration dropped a column billing reads.

3. Specs remove the variance, gaps included

The same vague prompt produced two coupon policies and three ways to handle a lapsed coupon. The spec runs agreed with each other every time, including on a case the spec forgot, which one run pointed out.

The three cases

CaseRunsPassed every hidden checkWhat went wrong without a spec
Checkout coupon codes6Vague 2/3 · Spec 3/3One run allowed unlimited reuse of a coupon; the three runs split three ways on charging a lapsed coupon.
API error envelope9No docs 0/3 · Docs 0/3 · Spec 3/3Without the contract doc, 2 of 3 runs changed body.error to an object and broke mobile v3.x. With it, all kept two error shapes.
Split-name migration9No rules 2/3 · Rules 3/3 · Spec 3/3One run dropped users.name and changed createUser's signature. All six non-spec runs split names at the first space.

Checkout coupon codes

"Add coupon codes to checkout." Two of three runs got everything right. The third decided coupons have no usage limit, and said so.

Open the coupon runs

API error envelope

"Clean up our API errors." A well-designed envelope that shows every mobile user a blank error banner, and the 13-line doc that prevented it.

Open the API runs

Split-name migration

"Split name into first_name and last_name." A correct backfill followed by DROP COLUMN name, with a warning in the summary nobody would act on.

Open the migration runs

New: the same runs on four Claude models

We re-ran all three cases on Sonnet 5.5, Haiku 4.5 and Fable 5.1, 78 runs in total. With a spec, 28 of 30 runs passed every hidden check across all four models. Without one, 12 of 48 did, and the model Anthropic ranks higher on Terminal-Bench 4.0 was the one that dropped a column billing reads in 3 of 3 runs.

Compare Opus 5.5, Sonnet 5.5, Haiku 4.5 and Fable 5.1

How the runs were recorded

Same conditions for every run

  • Claude Code 2.1.284, model claude-opus-5-5, headless claude -p, recorded September 28, 2026.
  • A fresh copy of the fixture repo per run. No user settings, hooks, plugins or MCP servers (--setting-sources project --strict-mcp-config).
  • One prompt, no follow-up messages, no human edits before scoring.
  • Runs took 37 to 89 seconds and 3 to 8 turns.

How scoring works

  • Each case has a hidden acceptance suite, written and frozen before the first run and never shown to the agent.
  • Suites call only interfaces that existed before the change (functions or HTTP), and accept any reasonable way of reporting an error, so a different API design is not a failure.
  • Where a check tests a rule the vague prompt never stated, the case page labels it.
  • One harness bug was found and fixed after the runs; the migration page explains it and the pack includes both suite versions.

Run packs

Each pack contains the fixture repo, the exact prompts, the hidden suite, the runner script, and for every run on every model its diff, hidden-suite result, own test counts and final message. Licensed CC BY 4.0.

What this does not show

Three runs per variant is enough to show that something can happen, not how often it happens. The repos are small and clean on purpose, so an outcome can be traced to one prompt or one file; larger codebases give an agent more to misread. We wrote the specs, the docs and the suites, which is how specs work in practice, but it means the specs cover what we test. The one gap a spec left open, in the coupon case, is reported rather than hidden.

If you repeat these with another model or tool, we would like to publish the results next to ours. Send them to us.

Turn a one-line ticket into the rules the agent will follow

Start with the spec packet generator, or read how spec-first work fits alongside agent harnesses.

Editorial note

Every number on these pages comes from the recorded runs in the run packs. We fixed nothing in the agents' output before scoring it.