Building an RL task for coding agents · case study

What Agents Don't Test

I built a task for coding agents three times. Each version taught me something about grading them, and all three pointed to the same failure. Three pilots at larger scale then showed where that failure doesn't appear.

Rushil Jaiswal · October 2026 · Opus 4.7, Opus 5.5 and Sonnet 5.5 · 33 valid graded runs · Code: what-agents-dont-test and scaling-attempt

In short: I built an RL task where a coding agent must reimplement a tool's behavior exactly, graded against a reference implementation on hundreds of hidden cases. Over three versions, the failing runs all made the same mistake: they tested only inside their own assumptions and reported that as proof they were done. When I tried a larger task in a real codebase, in three pilots frontier models handled readable migrations and known bug patterns easily, which sharpened the claim: failures come from behavior that has to be discovered, not from size. Next is putting that discovery problem inside a real codebase.

The finding: agents test extensively, but only inside their own assumptions, and then report that testing as proof they're done. The failing runs each said thousands of their own random tests matched the reference. Each one had missed exactly the behavior its tests never generated.

The task

A small Python tool, snapshot, copies a project and skips files listed in ignore files. It currently asks a reference program which files to skip. The agent is told that the reference won't exist where the tool is deployed, and must rewrite one function, ignored_files(root), in pure Python so it returns exactly what the reference returns, for any directory tree.

Grading is differential: I generate hundreds of hidden random trees, record the reference's answer for each, and count the cases where the agent's answer matches exactly. No answer is ever written by hand, so I can't bake my own misunderstanding into the grader. The reference stays available while the agent works, so a careful agent can test itself against it. Whether it actually does is part of what the task measures.

This is the same grading idea as Mechanize's GBA Eval, which compares an agent-written Game Boy Advance emulator against a reference emulator frame by frame, at a much smaller scale.

v1 · real git3/3

Opus 5.5 solved, ~12 min each, by recalling git's source

v2.1 · invented tool3/4

Opus 5.5 solved, ~30 min, by experimenting

v2.2 · non-path rules4/9

Opus 5.5 solved; usually fails

v1: gitignore, graded against real git

The first reference was git itself. The agent reimplements gitignore matching; the grader compares against git ls-files --ignored on random trees. To keep the expected answers from depending on my machine, the oracle runs git with a fake home directory, no global ignore file, case-sensitive matching, and an empty repo template, and refuses to answer if git's two lists (ignored and not ignored) don't partition the files exactly.

My first case generator used realistic names, and Opus 4.7's very first draft scored 100% on it. The agent's own fuzzer was better than mine: it used a tiny alphabet so patterns collided in odd ways. I added a hard profile built on that idea (glued and repeated **, POSIX classes, reversed ranges, escapes, CRLF), and the picture changed:

ModelRunsOriginal setHard setSolved
Opus 4.74100%83–91%0/4
Opus 5.53100%100%3/3

A minimizer that deletes ignore-file lines while the candidate still disagrees with git reduced every Opus 4.7 failure to a root cause. Two were in all four runs: reversed ranges like [b-a] crashed its regex translation, and POSIX classes like [[:alpha:]] weren't supported. Every run had fuzzed against git and reported a match. None of their fuzzers generated those inputs.

Opus 5.5 solved v1 in about 12 minutes per run. Its transcripts explain why: it ported git's wildmatch.c nearly line for line from memory. A task built on a famous specification tests recall, not engineering.

v2: an invented tool behind a black box

v2 replaces git with packignore, a legacy tool I invented, so there's nothing to recall. Its source never enters the agent's sandbox: it runs as a server on the host, and the agent gets a thin client that uploads a tree and prints the answer. The agent also gets an old README that is incomplete, and wrong in one place. Five rules are documented; the rest the agent has to discover:

RuleTruthWhy it's a test
W1Comments start with ;. The README says #.Trust the docs or check them?
H1The most specific pattern wins, not the last oneContradicts the gitignore habit every model has
H2** spans at most 3 directoriesOnly visible in deep trees
H3! can re-include files inside an excluded directoryHalf like git, half not
H4{a,b} alternation worksSyntax absent from the docs

Each hidden rule has a reference test showing that a 2–3 line experiment reveals it. Every hidden case is tagged by which rules it exercises (switch a rule off; if the answer changes, the case tests that rule), and the hidden set holds 100 cases per rule, so no single quirk decides the score.

The comment rule needed evidence

In the first two clean runs, Opus 5.5 found every rule except one. All four Opus 5.5 runs, including two pilots, tested that # isn't a comment, concluded there are no comments, and never tried another character. Nothing in the environment pointed at ;, so the only route was guessing. That's the grader checking something the agent had no reason to look for. I added a legacy .packignore to the starter repo that uses ; comments (v2.1). After that, Opus 5.5 solved 3 of 4 valid runs in 27–42 minutes, and Opus 4.7 scored 43% and 66%, running out of its $10 budget both times.

v2.2: rules that don't depend on the path

Every Opus 5.5 transcript discovered rules by varying paths and patterns. None varied anything else about a file. v2.2 adds two rules that ignore the path entirely: a file whose first line contains @generated is excluded by default, and so is any executable file. A ! pattern can re-include either.

The evidence sits in the agent's own repo. It contains a generated version file and two executable scripts, and running the tool on the repo itself excludes all three even though no pattern matches them. To keep the client from giving the game away, it uploads the whole tree as a tar archive, which carries contents and permissions without saying which ones matter.

LimitRunsSolvedSolve timeHow failures ended
60 min / $106226, 51 min2 still searching at the limit; 2 stopped early, confident
2 h / $203232, 33 min1 stopped itself with 75 minutes left
All9426–51 min3 of 5 failures: confident and wrong
Score over time, solved v2.2 run: 73.9% at 8 minutes, 84.6% at 37 minutes, 100% at 51 minutes 100% 75% 50% 0 15 30 45 60 min time limit min 8: 73.9%, both new rules found min 37: 84.6% min 51: 100%
One solved v2.2 run, rebuilt from its transcript's edits and scored version by version. It found the two new rules early and spent most of the hour on the rest.

The pattern across all three versions

Every failing run built a fuzzer, compared against the reference, and reported success. What it couldn't do was generate inputs outside its own model of the tool.

WhereMissedHow oftenWhat its own testing said
v1, Opus 4.7POSIX classes, reversed ranges4 of 4 runsMatched git
v2.0, Opus 5.5; comments4 of 4 runs"No comments" after ruling out #
v2.2, Opus 5.5Braces, generated and executable files, the depth limit3 of 5 failures~7,000, ~9,000 and ~4,500 random trees "identical"
+ every v2.2 run read the generated file and the release script in its first minute - two independent runs wrote different code (220 and 150 lines) with identical answers on all 800 hidden cases: the same wrong theory - one run confirmed both new rules at minute 59, one minute before the limit - one 2-hour run declared itself done at minute 45, missing the depth limit

Extra time helps an agent that is still searching. It doesn't help one that believes it's finished. Doubling the limit to two hours didn't change the solve times (both 2-hour solves finished in about 33 minutes) and didn't rescue the run that stopped early.

What I got wrong along the way

More than twenty problems in my own grader and harness were caught and fixed before they could distort a result. These were the instructive ones:

I also made wrong calls in the analysis and corrected them on the record: I first judged brace expansion safe, first said the time limit didn't matter for a run that found the missing rules at minute 59, and first blamed a slowdown on an agent's probes when replaying them showed they were cheap.

Trying to scale it up: three pilots

The obvious weakness of this task is scale: one file, about 30 minutes. So I tried to build a second task inside a real, sizable codebase, and before building any grading infrastructure I ran cheap pilots to check that the failure I wanted to measure actually exists. It didn't, three times.

The codebase was Pelican, a static site generator (about 8,700 lines of source), chosen from ten measured candidates because it has a deterministic boundary to grade against (the generated site), reference-output tests already in its suite, and no prior upstream work on the changes I tried.

PilotTaskRunsResult
1Migrate the source from os.path to pathlib without changing behavior2Both passed every check in ~10 min, by keeping os.path exactly where pathlib behaves differently
2Same, but filesystem paths must use pathlib (no sidestepping the risky spots)4 (3 valid)Every valid run passed, across three models; it separated them on cost (Sonnet 5.5 $1.80, Opus 4.7 $19.37), not correctness
3Fix a real bug I found during pilot 1: with some plugins, output differs between builds (since submitted upstream with a fix and regression test: pelican#3621)4All four fixed the root cause in 1–5 minutes

Each pilot had hidden checks the agent couldn't see, and I validated each check with a control that should fail it before trusting a pass. For the migration: a site with relative URLs that the visible tests never exercise, a probe plugin checking that values reaching plugins are still strings, a theme template doing string operations on paths, incremental rebuilds, a cache written by the old code and read by the new, and unusual filesystem states. For the bug: determinism across hash seeds, determinism when Pelican runs as a library rather than from the command line, unchanged output for sites that were already reproducible, and an unchanged public API. Two tempting wrong fixes I prepared for (changing a return type, forcing a fixed hash seed) both passed Pelican's own tests and both failed exactly one hidden check. No agent tried either.

What the pilots taught me

About $70 of agent time answered that before I built a grader. Every run, check and decision, with the commit history: github.com/rushjais/scaling-attempt.

Limits

What I'd build next

Built with Claude Code as a coding agent under my direction: I set the direction, made the design and scoping calls, reviewed results and pushed back, and the code and much of the prose were written by the agent. Every number on this page comes from graded runs in the repository, including the invalid ones, which are kept and labeled. Code and run logs: github.com/rushjais/what-agents-dont-test (this task) and github.com/rushjais/scaling-attempt (the pilots). Related: vaudit, on whether graders are fair to correct work.