Checkout coupon codes: one-line prompt vs spec, six real runs

We gave Claude Code the same small checkout repo six times. Three runs got a one-line ticket, three got a spec. A hidden test suite, written before any run, scored the results. The vague prompt mostly produced working code. What it did not produce was the same product twice.

Vague prompt2 of 3 runs passed all 11 checks
Spec3 of 3 runs passed all 11 checks
The real gapVague runs chose 2 coupon limits and 3 payment behaviours

The setup

The repo

A Node checkout service with no dependencies: totals.js owns the money math, store.js holds products, five coupons and an empty redemptions list, and applyCoupon() is a stub that throws. Two existing tests pass.

Seeded coupons cover the cases that usually go wrong: 20% off, $10 off, $50 off (bigger than some carts), one expired, one disabled.

The runs

Claude Code 2.1.284, claude-opus-5-5, headless (claude -p), a fresh copy of the repo each time, no user settings, hooks or MCP servers. Each run took 45 to 66 seconds and 3 or 4 turns. Recorded on September 28, 2026.

Every file is in the run pack: fixture, prompts, hidden suite, and each run's diff and final message.

The two prompts, verbatim

Vague (runs V1 to V3)

Add coupon codes to checkout. Make sure invalid codes do not break payment.

Spec (runs S1 to S3), abridged

Non-goals
- No admin, no stacking, no tax or rounding changes.
- Do not modify src/totals.js or src/processor.js.

Rules
- A coupon can be redeemed once per user. It counts as
  redeemed only when an order is paid.
- Applying a second coupon replaces the first.

Acceptance criteria
- AC-2 Unknown, disabled or expired code: applyCoupon
  returns { ok: false, reason }. It does not throw.
- AC-4 Redemption is re-checked inside pay.
- AC-5 Amount-off larger than the subtotal floors at 0.
  ... six ACs in total

Hidden suite scorecard

The suite only calls functions that already existed in the repo, and it accepts a rejection whether it arrives as a thrown error or as an error result, so it does not punish the vague runs for choosing a different API. Checks marked unstated test rules the one-line ticket never mentions.

CheckKindVagueSpec
H1 Percent coupon discounts the pre-tax subtotalstated3/33/3
H2 Processor is charged exactly the displayed totalstated3/33/3
H3 Expired code leaves the total unchanged, payment worksstated3/33/3
H4 Disabled code is not appliedstated3/33/3
H5 Unknown code leaves the checkout payablestated3/33/3
H6 A user cannot redeem the same coupon twiceunstated2/33/3
H7 One user's redemption does not block another userunstated3/33/3
H8 Applying without paying does not use the coupon upunstated3/33/3
H9 A second coupon replaces the first, no stackingunstated3/33/3
H10 Amount-off larger than the cart never charges below 0correctness3/33/3
H11 Two open checkouts cannot both spend a one-per-user couponunstated2/33/3

On the untouched repo the "reject" checks (H3 to H6, H10, H11) pass trivially because no coupon ever applies, so they only mean something alongside H1 and H2. Every run passed its own tests: the vague runs wrote 10 to 12 new tests, the spec runs 10 or 11.

Where the runs actually differed: decisions

The scorecard is close. The decisions are not. We measured these by calling each run's code directly, not by reading its summary.

DecisionV1V2V3S1 to S3
How often can a user use a coupon?OnceUnlimitedOnceOnce (spec)
Coupon expires between apply and payCharges full price, $54.00Refuses the first pay()Refuses the first pay()Charges the discounted $43.20
How a bad code is reportedThrows CouponErrorThrowsThrowsReturns { ok: false, reason }
Code matchingCase-insensitiveCase-insensitiveCase-insensitiveExact
New public APIremoveCouponremoveCouponremoveCoupon, new coupons.jsNone

Same prompt, two coupon policies

V1 read the per-user redemptions table as a hint and enforced one use per user. V2 read the same data and concluded: "the coupon data has no per-user or single-use limit, so none is enforced." Both said so in their final message. Only one matches what the business wanted.

Three answers to one money question

When a coupon lapses before payment, V1 quietly charges more than the customer was shown; V2 and V3 refuse the payment once, which the frontend now has to handle. Each is defensible. None was a choice anyone on the team made.

The frontend contract drifts too

Every vague run throws on a bad code; the spec runs return { ok, reason }. A frontend written for one gets an uncaught exception, or silently ignores the error, on the other. The ticket never said which contract the frontend team was building against.

What the spec did not fix

All three spec runs honoured a coupon that expired after it was applied, and charged the discounted total. The spec told pay to re-check redemption and said nothing about expiry, so the agent did exactly that, three times out of three. S1 flagged it in its final message:

Re-checks at pay: pay only re-checks redemption, as AC-4 asks. A coupon that
expires or is disabled after being applied is still honoured at payment. If
you'd rather reject it there, pay can reuse the same checks applyCoupon runs.

This is the honest shape of the result. A spec does not make the agent smarter; it makes the agent's behaviour predictable, gaps included. The gap here was visible in one line of the spec and one line of the run summary, instead of being spread across three different implementations.

The line that decides H11

Two open checkouts, one user, one single-use coupon. Whether both get the discount depends on one re-check at payment time. From run S1, src/checkout.js:

export function pay(checkoutId) {
  const co = mustGet(checkoutId);
  if (co.status !== 'open') throw new Error(`checkout ${checkoutId} is ${co.status}`);
  // The coupon may have been redeemed on another order since it was applied here.
  if (co.couponCode && isRedeemed(co.userId, co.couponCode)) co.couponCode = null;
  const summary = getSummary(checkoutId);
  const ch = charge(co.userId, summary.total);
  ...
  if (co.couponCode) db.redemptions.push({ code: co.couponCode, userId: co.userId, orderId: order.id });

V1 and V3 wrote an equivalent check without being asked. V2 had no per-user rule at all, so it had nothing to re-check.

What to take from six runs

Do not review for "does it work"

On a small, well-structured repo, a current model usually ships working code from a one-liner. Review for which rules it picked: limits, what happens to money on the unhappy path, and the error contract.

Write down the rules that cost money

The two lines that removed most of the variance were "once per user, counted at payment" and "bad codes return a result, never throw". Those belong in the spec even when everything else is left to the agent.

Read the final message as a diff

Every run listed its assumptions. The vague runs' assumptions are the review checklist the ticket should have contained; turn them into acceptance criteria before merge.

Limits of this test

  • Three runs per variant is enough to show variance, not to estimate a rate. "2 of 3" means "happened once", not "fails a third of the time".
  • The repo is small and clean. Larger codebases give the agent more to misread and more places to put a coupon rule.
  • We wrote both the spec and the hidden suite, so the spec naturally covers what we test. The suite was frozen before the first run and never shown to the agent; the spec's expiry gap shows it was not written to the test.
  • One model and one tool. Results for other agents may differ; the run pack is there so you can check.

The same case on Sonnet 5.5, Haiku 4.5 and Fable 5.1

Without a spec, Haiku 4.5 failed both coupon-reuse checks in 3 of 3 runs; with one it passed all 11 checks every time.

See the cross-model results

Related

The full coupon spec packet

The same feature as a reviewer-ready packet: tasks, acceptance criteria, QA fixtures and rollback.

Open the coupon case

API error envelope runs

Nine runs where one missing doc decided whether the mobile app broke.

Open the API runs

Split-name migration runs

Nine runs on a schema change, including one that dropped a column other services read.

Open the migration runs

Write the rules before the agent picks them

The generator turns a one-line ticket into non-goals, rules and acceptance criteria you can paste above the prompt.

Editorial note

Every number on this page comes from the recorded runs in the run pack. We fixed nothing in the agents' output before scoring it.