01 The task
A small tool decides which files to skip by calling a reference program. The agent has to rewrite that function in pure Python so it returns exactly what the reference returns, for any directory tree.
Grading is differential: hundreds of hidden random trees, the reference’s answer recorded for each, and an exact match required. No expected answer is written by hand, so the grader can’t encode my own misunderstanding. The reference stays available while the agent works, so a careful agent can test itself against it. Whether it does is part of what the task measures.
02 Three versions, each answering the last one’s weakness
| Version | Reference | Opus 5.5 | What it taught me |
|---|---|---|---|
| v1 | real git (gitignore) | 3 of 3 solved, ~12 min | It ported git’s wildmatch.c from memory. A famous spec tests recall. |
| v2.1 | an invented tool behind a black box | 3 of 4 solved, ~30 min | Nothing to recall, so it had to experiment. It separated models; Opus 4.7 scored 43–66%. |
| v2.2 | + rules that ignore the path (file contents, permissions) | 4 of 9 solved | The first version the newest model usually fails, for a reason I can explain. |
03 The pattern across all three
Each failing run tested extensively, but only inside its own picture of the tool. In v2.2, three runs reported that roughly 7,000, 9,000 and 4,500 of their own random trees matched the reference, and scored 52%, 68% and 87% on the hidden set. Two independent runs wrote different code with identical answers on all 800 hidden cases: the same wrong theory. Every run read the evidence in its first minute; what separated success from failure was whether it followed up.
Giving a run twice the time didn’t change that. The 2-hour solves finished in about 33 minutes, and the 2-hour failure declared itself done with 75 minutes left.
04 Trying to scale it up
To answer the obvious objection (one file, 30 minutes) I tried a second task inside a real codebase, Pelican, and ran cheap pilots before building any infrastructure: a behavior-preserving os.path→pathlib migration, a stricter version of it, and a real bug I found along the way. Every hidden check was validated with a control that should fail it.
All three pilots came back clean. Every valid run passed, three models separated only on cost, and the bug was fixed by all four runs in 1–5 minutes. That sharpened the claim: in these pilots, when the specification was readable, careful agents didn’t miss it. The failures came from behavior that had to be discovered, not from size.
Honest limits. The task is small: one file, about 30 minutes per run, far from a 24-hour emulator build. The tool in v2 is invented, which defeats recall but can read as contrived. Samples are 2 to 9 valid runs per version or pilot, three models from one provider, one harness. More than twenty bugs in my own grader and harness were caught along the way, and invalid runs are kept and labeled rather than counted.