1. Vague prompts mostly work
On small, readable repos the model fixed crashes, closed stack leaks, capped discounts and backfilled data correctly from one sentence. If your review question is "does it run", a current model often passes it.
We gave Claude Code three ordinary tickets, a coupon feature, an error-format cleanup and a schema change, and ran each one several times: from a one-line prompt, from a spec, and in some cases with the repo's written rules removed. Hidden test suites, frozen before the first run, scored every result. Every prompt, diff and score is downloadable.
On small, readable repos the model fixed crashes, closed stack leaks, capped discounts and backfilled data correctly from one sentence. If your review question is "does it run", a current model often passes it.
A 13-line mobile contract doc and three README lines were enough to stop the breaking changes, even with a one-line prompt. Without them, two of three API runs broke the shipped app and one migration dropped a column billing reads.
The same vague prompt produced two coupon policies and three ways to handle a lapsed coupon. The spec runs agreed with each other every time, including on a case the spec forgot, which one run pointed out.
| Case | Runs | Passed every hidden check | What went wrong without a spec |
|---|---|---|---|
| Checkout coupon codes | 6 | Vague 2/3 · Spec 3/3 | One run allowed unlimited reuse of a coupon; the three runs split three ways on charging a lapsed coupon. |
| API error envelope | 9 | No docs 0/3 · Docs 0/3 · Spec 3/3 | Without the contract doc, 2 of 3 runs changed body.error to an object and broke mobile v3.x. With it, all kept two error shapes. |
| Split-name migration | 9 | No rules 2/3 · Rules 3/3 · Spec 3/3 | One run dropped users.name and changed createUser's signature. All six non-spec runs split names at the first space. |
"Add coupon codes to checkout." Two of three runs got everything right. The third decided coupons have no usage limit, and said so.
Open the coupon runs"Clean up our API errors." A well-designed envelope that shows every mobile user a blank error banner, and the 13-line doc that prevented it.
Open the API runs"Split name into first_name and last_name." A correct backfill followed by DROP COLUMN name, with a warning in the summary nobody would act on.
We re-ran all three cases on Sonnet 5.5, Haiku 4.5 and Fable 5.1, 78 runs in total. With a spec, 28 of 30 runs passed every hidden check across all four models. Without one, 12 of 48 did, and the model Anthropic ranks higher on Terminal-Bench 4.0 was the one that dropped a column billing reads in 3 of 3 runs.
Compare Opus 5.5, Sonnet 5.5, Haiku 4.5 and Fable 5.1claude-opus-5-5, headless claude -p, recorded September 28, 2026.--setting-sources project --strict-mcp-config).Each pack contains the fixture repo, the exact prompts, the hidden suite, the runner script, and for every run on every model its diff, hidden-suite result, own test counts and final message. Licensed CC BY 4.0.
Three runs per variant is enough to show that something can happen, not how often it happens. The repos are small and clean on purpose, so an outcome can be traced to one prompt or one file; larger codebases give an agent more to misread. We wrote the specs, the docs and the suites, which is how specs work in practice, but it means the specs cover what we test. The one gap a spec left open, in the coupon case, is reported rather than hidden.
If you repeat these with another model or tool, we would like to publish the results next to ours. Send them to us.
Start with the spec packet generator, or read how spec-first work fits alongside agent harnesses.
Every number on these pages comes from the recorded runs in the run packs. We fixed nothing in the agents' output before scoring it.