work / what agents don’t test

What agents don’t test

Every failing agent built a fuzzer, compared against the reference, and reported a match. Each had missed exactly what its fuzzer never generated.

An RL task for coding agents, built three times and graded against a reference implementation. Opus 5.5 recalled the first version from memory; the third it solved in 4 of 9 runs.

Status
Research · continues vaudit from auditing graders to building a task
Role
Sole author, directing a coding agent. 33 valid graded runs across Opus 4.7, Opus 5.5 and Sonnet 5.5
Stack
Python · Docker · Claude Code (headless) · differential testing
Write-up
The full write-up, with every run →
Source
what-agents-dont-test ↗ · scaling-attempt ↗
Upstream
A bug found along the way, fixed with a regression test: pelican#3621 ↗

01  The task

A small tool decides which files to skip by calling a reference program. The agent has to rewrite that function in pure Python so it returns exactly what the reference returns, for any directory tree.

Grading is differential: hundreds of hidden random trees, the reference’s answer recorded for each, and an exact match required. No expected answer is written by hand, so the grader can’t encode my own misunderstanding. The reference stays available while the agent works, so a careful agent can test itself against it. Whether it does is part of what the task measures.

02  Three versions, each answering the last one’s weakness

VersionReferenceOpus 5.5What it taught me
v1real git (gitignore)3 of 3 solved, ~12 minIt ported git’s wildmatch.c from memory. A famous spec tests recall.
v2.1an invented tool behind a black box3 of 4 solved, ~30 minNothing to recall, so it had to experiment. It separated models; Opus 4.7 scored 43–66%.
v2.2+ rules that ignore the path (file contents, permissions)4 of 9 solvedThe first version the newest model usually fails, for a reason I can explain.

03  The pattern across all three

Each failing run tested extensively, but only inside its own picture of the tool. In v2.2, three runs reported that roughly 7,000, 9,000 and 4,500 of their own random trees matched the reference, and scored 52%, 68% and 87% on the hidden set. Two independent runs wrote different code with identical answers on all 800 hidden cases: the same wrong theory. Every run read the evidence in its first minute; what separated success from failure was whether it followed up.

Giving a run twice the time didn’t change that. The 2-hour solves finished in about 33 minutes, and the 2-hour failure declared itself done with 75 minutes left.

04  Trying to scale it up

To answer the obvious objection (one file, 30 minutes) I tried a second task inside a real codebase, Pelican, and ran cheap pilots before building any infrastructure: a behavior-preserving os.path→pathlib migration, a stricter version of it, and a real bug I found along the way. Every hidden check was validated with a control that should fail it.

All three pilots came back clean. Every valid run passed, three models separated only on cost, and the bug was fixed by all four runs in 1–5 minutes. That sharpened the claim: in these pilots, when the specification was readable, careful agents didn’t miss it. The failures came from behavior that had to be discovered, not from size.

Honest limits. The task is small: one file, about 30 minutes per run, far from a 24-hour emulator build. The tool in v2 is invented, which defeats recall but can read as contrived. Samples are 2 to 9 valid runs per version or pilot, three models from one provider, one harness. More than twenty bugs in my own grader and harness were caught along the way, and invalid runs are kept and labeled rather than counted.