Splitting a name column with an AI agent: nine runs, one dropped column

"Split the users' name field into first_name and last_name" sounds like a ten-minute ticket. It is also a schema change that other systems depend on. We ran it nine times in Claude Code against a small SQLite service, then ran every result against a production-like database with awkward real names.

No repo rules1 of 3 dropped users.name
Rules in README3 of 3 passed all 8 checks
Spec3 of 3 passed, and split names the way the team asked

The setup

The repo

An accounts service on node:sqlite with a tiny migration runner, two applied migrations, and createUser, renameUser, getUser and exportCsv. Three tests pass.

The README states three rules: applied migrations must never be edited; a separate billing sync reads users.name straight from the database; the CRM imports exportCsv and matches on its exact header.

Three variants, three runs each

  • No rules (N1 to N3): vague prompt, README rules and the CRM comment removed.
  • Rules (V1 to V3): vague prompt, README rules present.
  • Spec (S1 to S3): a spec prompt, README rules present.

Claude Code 2.1.284, claude-opus-5-5, headless, fresh copy each run, 42 to 89 seconds each. Recorded September 28, 2026. Run pack.

The two prompts

Vague (N and V runs), verbatim

Split the users' name field into first_name and last_name.

Spec (S runs), abridged

Goal: expand step of expand/contract. Dropping name
is a later, separate change.

Non-goals
- Do not drop, rename or stop writing users.name.
- Do not change exportCsv output or its header.

Split rule
- Collapse whitespace, trim. One token → first_name,
  last_name ''. Otherwise last token → last_name.
- NULL name → both new columns NULL.

AC-1 New migrations/003_*.sql works on a DB with rows.
AC-3 createUser({ email, name }) keeps working and
     fills all three columns; renameUser keeps them in sync.

Hidden suite scorecard

The suite builds a database the way production has it: migrations 001 and 002 applied, then seven existing users including Cher, Mary Ann Smith, Smith, Jr., a NULL name and one with stray whitespace. Then it runs each build's new migrations and checks what survived. It does not score how names are split, only that nothing is lost.

CheckNo rulesRulesSpec
D1 New migrations apply to a database that already has rows3/33/33/3
D2 Applied migrations 001 and 002 were not edited3/33/33/3
D3 Backfill loses no characters of any name3/33/33/3
D4 users.name is still there, unchanged, for the billing sync2/33/33/3
D5 CSV export for existing users is byte-for-byte unchanged2/33/33/3
D6 Running migrate again is a no-op3/33/33/3
D7 Existing createUser({ email, name }) callers still work2/33/33/3
D8 renameUser keeps all name columns in sync2/33/33/3

Harness note: our first scoring pass computed the "before" CSV with each build's own exportCsv, which crashed on N3 and failed D1 to D8 for the wrong reason. We fixed the harness to use the original function and re-scored all nine runs from their saved output; no agent was re-run. Both versions of the suite are in the run pack.

Run N3: a clean migration that breaks billing

-- migrations/003_split_users_name.sql (run N3)
ALTER TABLE users ADD COLUMN first_name TEXT;
ALTER TABLE users ADD COLUMN last_name TEXT;

UPDATE users SET
  first_name = CASE WHEN instr(trim(name), ' ') > 0
    THEN substr(trim(name), 1, instr(trim(name), ' ') - 1)
    ELSE nullif(trim(name), '') END,
  last_name = CASE WHEN instr(trim(name), ' ') > 0
    THEN trim(substr(trim(name), instr(trim(name), ' ') + 1)) END
WHERE name IS NOT NULL;

ALTER TABLE users DROP COLUMN name;

The backfill is correct and loses nothing. Then the last line removes the column the billing sync reads, and createUser and renameUser switch to { firstName, lastName }, so every existing caller breaks too. The agent was not careless. Its final message said:

It then drops the name column. This can't be undone by re-running migrations.
If anything outside this code reads users.name directly from the database,
remove the DROP COLUMN line and drop the column later.

That warning is the whole case for writing constraints down. The agent knew a reader might exist and could not know one did. N1 and N2, with the same prompt and the same missing README, kept the column. A reviewer who skims the summary gets a migration that passes every test in the repo and fails in production the next night.

Decisions nobody asked for

Measured by migrating the same four names through every build.

Input nameAll six non-spec runsS1 to S3
Mary Ann SmithMary / Ann SmithMary Ann / Smith
Ludwig van BeethovenLudwig / van BeethovenLudwig van / Beethoven
CherCher / NULLCher / ''
Smith, Jr.Smith, / Jr.Smith, / Jr.

Consistent, and still a choice

Every non-spec run split at the first space. Neither rule is right for every name; the point is that one of them was chosen by the model, in six out of six runs, and would have become the data in every downstream system.

Both rules fail "Smith, Jr."

No run handled suffixes, and the spec did not ask. If your data has them, that is a line for the spec, not something to hope the agent notices.

Tests pass either way

Every run tested its backfill on inserted rows, and every run's own tests passed, including N3's. Spec runs added 6 to 8 tests, the others 1 to 3. None of those tests could see the billing sync, because nothing in the repo does.

What to take from nine runs

Name your readers

Three lines of README ("billing reads users.name", "CRM matches the CSV header", "never edit applied migrations") took the vague prompt from 1 in 3 dangerous to 3 of 3 safe. Most schema breakages are readers outside the repo.

Say "expand only"

An agent asked to "split" a column will often finish the job and remove the old one. If the change is expand/contract, say which step this is.

Test on data, not on an empty schema

A migration that passes on a fresh database says little. The check that matters seeds the old schema with real-shaped rows first.

Limits of this test

  • Three runs per variant shows what can happen, not how often. N3 is one run.
  • The README rules were short and prominent. In a large repo the same sentence can sit in a doc the agent never opens.
  • We wrote the rules, the spec and the hidden suite. The suite's checks were fixed before the first run and never shown to the agent; the harness fix above changed how the baseline was computed, not what was checked.
  • One model and one tool. The run pack has everything needed to repeat this.

The same case on Sonnet 5.5, Haiku 4.5 and Fable 5.1

With no written rules, Sonnet 5.5 dropped users.name in 3 of 3 runs. With the README rules present, Haiku 4.5 still dropped it in 3 of 3.

See the cross-model results

Related

The database schema spec packet

Migration plan, backfill checks and rollback for a reviewer-ready schema change.

Open the database case

API error envelope runs

Nine runs where one missing doc decided whether the mobile app broke.

Open the API runs

Checkout coupon runs

Six runs where the vague prompt worked but picked different business rules each time.

Open the coupon runs

Spec the migration before the agent writes it

The database spec generator walks through readers, backfill, expand/contract steps and rollback.

Editorial note

Every number on this page comes from the recorded runs in the run pack. We fixed nothing in the agents' output before scoring it.