What Agents Don't Test
I built a task for coding agents three times. Each version taught me something about grading them, and all three pointed to the same failure. Three pilots at larger scale then showed where that failure doesn't appear.
In short: I built an RL task where a coding agent must reimplement a tool's behavior exactly, graded against a reference implementation on hundreds of hidden cases. Over three versions, the failing runs all made the same mistake: they tested only inside their own assumptions and reported that as proof they were done. When I tried a larger task in a real codebase, in three pilots frontier models handled readable migrations and known bug patterns easily, which sharpened the claim: failures come from behavior that has to be discovered, not from size. Next is putting that discovery problem inside a real codebase.
The finding: agents test extensively, but only inside their own assumptions, and then report that testing as proof they're done. The failing runs each said thousands of their own random tests matched the reference. Each one had missed exactly the behavior its tests never generated.
The task
A small Python tool, snapshot, copies a project and skips files listed in ignore files. It currently asks a reference program which files to skip. The agent is told that the reference won't exist where the tool is deployed, and must rewrite one function, ignored_files(root), in pure Python so it returns exactly what the reference returns, for any directory tree.
Grading is differential: I generate hundreds of hidden random trees, record the reference's answer for each, and count the cases where the agent's answer matches exactly. No answer is ever written by hand, so I can't bake my own misunderstanding into the grader. The reference stays available while the agent works, so a careful agent can test itself against it. Whether it actually does is part of what the task measures.
This is the same grading idea as Mechanize's GBA Eval, which compares an agent-written Game Boy Advance emulator against a reference emulator frame by frame, at a much smaller scale.
Opus 5.5 solved, ~12 min each, by recalling git's source
Opus 5.5 solved, ~30 min, by experimenting
Opus 5.5 solved; usually fails
v1: gitignore, graded against real git
The first reference was git itself. The agent reimplements gitignore matching; the grader compares against git ls-files --ignored on random trees. To keep the expected answers from depending on my machine, the oracle runs git with a fake home directory, no global ignore file, case-sensitive matching, and an empty repo template, and refuses to answer if git's two lists (ignored and not ignored) don't partition the files exactly.
My first case generator used realistic names, and Opus 4.7's very first draft scored 100% on it. The agent's own fuzzer was better than mine: it used a tiny alphabet so patterns collided in odd ways. I added a hard profile built on that idea (glued and repeated **, POSIX classes, reversed ranges, escapes, CRLF), and the picture changed:
| Model | Runs | Original set | Hard set | Solved |
|---|---|---|---|---|
| Opus 4.7 | 4 | 100% | 83–91% | 0/4 |
| Opus 5.5 | 3 | 100% | 100% | 3/3 |
A minimizer that deletes ignore-file lines while the candidate still disagrees with git reduced every Opus 4.7 failure to a root cause. Two were in all four runs: reversed ranges like [b-a] crashed its regex translation, and POSIX classes like [[:alpha:]] weren't supported. Every run had fuzzed against git and reported a match. None of their fuzzers generated those inputs.
Opus 5.5 solved v1 in about 12 minutes per run. Its transcripts explain why: it ported git's wildmatch.c nearly line for line from memory. A task built on a famous specification tests recall, not engineering.
v2: an invented tool behind a black box
v2 replaces git with packignore, a legacy tool I invented, so there's nothing to recall. Its source never enters the agent's sandbox: it runs as a server on the host, and the agent gets a thin client that uploads a tree and prints the answer. The agent also gets an old README that is incomplete, and wrong in one place. Five rules are documented; the rest the agent has to discover:
| Rule | Truth | Why it's a test |
|---|---|---|
W1 | Comments start with ;. The README says #. | Trust the docs or check them? |
H1 | The most specific pattern wins, not the last one | Contradicts the gitignore habit every model has |
H2 | ** spans at most 3 directories | Only visible in deep trees |
H3 | ! can re-include files inside an excluded directory | Half like git, half not |
H4 | {a,b} alternation works | Syntax absent from the docs |
Each hidden rule has a reference test showing that a 2–3 line experiment reveals it. Every hidden case is tagged by which rules it exercises (switch a rule off; if the answer changes, the case tests that rule), and the hidden set holds 100 cases per rule, so no single quirk decides the score.
The comment rule needed evidence
In the first two clean runs, Opus 5.5 found every rule except one. All four Opus 5.5 runs, including two pilots, tested that # isn't a comment, concluded there are no comments, and never tried another character. Nothing in the environment pointed at ;, so the only route was guessing. That's the grader checking something the agent had no reason to look for. I added a legacy .packignore to the starter repo that uses ; comments (v2.1). After that, Opus 5.5 solved 3 of 4 valid runs in 27–42 minutes, and Opus 4.7 scored 43% and 66%, running out of its $10 budget both times.
v2.2: rules that don't depend on the path
Every Opus 5.5 transcript discovered rules by varying paths and patterns. None varied anything else about a file. v2.2 adds two rules that ignore the path entirely: a file whose first line contains @generated is excluded by default, and so is any executable file. A ! pattern can re-include either.
The evidence sits in the agent's own repo. It contains a generated version file and two executable scripts, and running the tool on the repo itself excludes all three even though no pattern matches them. To keep the client from giving the game away, it uploads the whole tree as a tar archive, which carries contents and permissions without saying which ones matter.
| Limit | Runs | Solved | Solve time | How failures ended |
|---|---|---|---|---|
| 60 min / $10 | 6 | 2 | 26, 51 min | 2 still searching at the limit; 2 stopped early, confident |
| 2 h / $20 | 3 | 2 | 32, 33 min | 1 stopped itself with 75 minutes left |
| All | 9 | 4 | 26–51 min | 3 of 5 failures: confident and wrong |
The pattern across all three versions
Every failing run built a fuzzer, compared against the reference, and reported success. What it couldn't do was generate inputs outside its own model of the tool.
| Where | Missed | How often | What its own testing said |
|---|---|---|---|
| v1, Opus 4.7 | POSIX classes, reversed ranges | 4 of 4 runs | Matched git |
| v2.0, Opus 5.5 | ; comments | 4 of 4 runs | "No comments" after ruling out # |
| v2.2, Opus 5.5 | Braces, generated and executable files, the depth limit | 3 of 5 failures | ~7,000, ~9,000 and ~4,500 random trees "identical" |
Extra time helps an agent that is still searching. It doesn't help one that believes it's finished. Doubling the limit to two hours didn't change the solve times (both 2-hour solves finished in about 33 minutes) and didn't rescue the run that stopped early.
What I got wrong along the way
More than twenty problems in my own grader and harness were caught and fixed before they could distort a result. These were the instructive ones:
- Machine-dependent answers. git defaults to case-insensitive matching on macOS, and my global ignore file leaked into the oracle. Either would have made the "right answer" depend on whose laptop ran the grader.
- Free reward. A candidate could just call git, or the starter could call the oracle with Python's absolute path. Candidates now run with an empty
PATH, the oracle is stopped before grading, and a static check flags subprocess and network imports. - The oracle froze under probing. An agent deliberately probed for limits with patterns like
a*a*a*…, and my regex-based matcher backtracked exponentially, freezing the server for the rest of the run. I replaced it with a linear-time matcher, verified identical on 250,000 random inputs and all 900 frozen cases. - The oracle leaked internals. One malformed request returned a Python traceback with a host path in it. Errors are now plain and generic.
- The harness ended runs early. In headless mode the session ends with the agent's turn, so a run that started fuzzing in the background and waited to be notified just stopped. The agent is now told it's running non-interactively.
- Laptop sleep. Runs stalled overnight and resumed hours later because the time limit used a clock that pauses during sleep. The runner now keeps the machine awake and enforces wall-clock time. Runs affected by sleep, restarts or harness bugs are marked invalid, not counted.
I also made wrong calls in the analysis and corrected them on the record: I first judged brace expansion safe, first said the time limit didn't matter for a run that found the missing rules at minute 59, and first blamed a slowdown on an agent's probes when replaying them showed they were cheap.
Trying to scale it up: three pilots
The obvious weakness of this task is scale: one file, about 30 minutes. So I tried to build a second task inside a real, sizable codebase, and before building any grading infrastructure I ran cheap pilots to check that the failure I wanted to measure actually exists. It didn't, three times.
The codebase was Pelican, a static site generator (about 8,700 lines of source), chosen from ten measured candidates because it has a deterministic boundary to grade against (the generated site), reference-output tests already in its suite, and no prior upstream work on the changes I tried.
| Pilot | Task | Runs | Result |
|---|---|---|---|
| 1 | Migrate the source from os.path to pathlib without changing behavior | 2 | Both passed every check in ~10 min, by keeping os.path exactly where pathlib behaves differently |
| 2 | Same, but filesystem paths must use pathlib (no sidestepping the risky spots) | 4 (3 valid) | Every valid run passed, across three models; it separated them on cost (Sonnet 5.5 $1.80, Opus 4.7 $19.37), not correctness |
| 3 | Fix a real bug I found during pilot 1: with some plugins, output differs between builds (since submitted upstream with a fix and regression test: pelican#3621) | 4 | All four fixed the root cause in 1–5 minutes |
Each pilot had hidden checks the agent couldn't see, and I validated each check with a control that should fail it before trusting a pass. For the migration: a site with relative URLs that the visible tests never exercise, a probe plugin checking that values reaching plugins are still strings, a theme template doing string operations on paths, incremental rebuilds, a cache written by the old code and read by the new, and unusual filesystem states. For the bug: determinism across hash seeds, determinism when Pelican runs as a library rather than from the command line, unchanged output for sites that were already reproducible, and an unchanged public API. Two tempting wrong fixes I prepared for (changing a return type, forcing a fixed hash seed) both passed Pelican's own tests and both failed exactly one hidden check. No agent tried either.
What the pilots taught me
- In these pilots, when the specification was readable, careful agents didn't miss it. In a behavior-preserving migration the old code is the spec, and every run built its own old-vs-new comparison without being asked (11, 21 and more site configurations). A bigger codebase adds work, not this kind of difficulty.
- Known bug patterns are solved instantly. "Sets iterate in hash order, so sort them" is something every model knows.
- The failures in the first task came from somewhere else: behavior that had to be discovered, outside anything the agent thought to test. That's the property a larger second task would need, not more lines of code.
- My checkers had bugs too. My probe plugin read URLs early, which in Pelican freezes where attachments go, so the checker changed the output it was observing and briefly made a correct run look broken. And one agent killed its own process: it terminated everything whose command line contained "pelican", and my runner had put the prompt, which starts with "Pelican…", in that command line.
About $70 of agent time answered that before I built a grader. Every run, check and decision, with the commit history: github.com/rushjais/scaling-attempt.
Limits
- Scale. The agent rewrites one file in about 30 minutes. Production tasks, like GBA Eval's 24-hour emulator, are far larger. My attempts at a larger task in a real codebase are above; they didn't produce failures.
- The tool is invented. That's deliberate, to defeat recall, but some will find it contrived.
- Small samples. 2 to 9 valid runs per version or pilot: enough for a direction, not a rate.
- Three models, all from one provider, run through one harness.
What I'd build next
- Discovery inside a real codebase. The pilots suggest the next task shouldn't be bigger, but should put the mechanism that worked here (behavior that has to be found by experiment, with evidence available) inside a real project: for example, a codebase that depends on a black-box service whose undocumented behavior matters. It would get its own pilot first, and longer runs in the cloud rather than on a laptop.
- Grading the report as well as the code. Every failing run here made a specific, checkable claim ("identical on 9,000 trees"). Scoring whether an agent's final claims are true would reward agents that know what their testing doesn't cover.
Built with Claude Code as a coding agent under my direction: I set the direction, made the design and scoping calls, reviewed results and pushed back, and the code and much of the prose were written by the agent. Every number on this page comes from graded runs in the repository, including the invalid ones, which are kept and labeled. Code and run logs: github.com/rushjais/what-agents-dont-test (this task) and github.com/rushjais/scaling-attempt (the pilots). Related: vaudit, on whether graders are fair to correct work.