01 The half nobody measures
Goodhart asked whether a grader can be gamed. There is a second question, and it is the one that costs real compute: is the grader fair to work that is actually correct?
Tooling exists for the first. For the second you need to know how often a grader rejects a legitimate solution, whether it agrees with itself on a rerun, and whether it asserts things the prompt never made knowable. Those numbers are not usually collected.
To measure them you need a task with a real, defensible disagreement in it, not a synthetic perturbation. CHIP-8 has five of them. The original COSMAC VIP and the later CHIP-48 read the same instructions differently, both defensibly, and real interpreters in the wild do both.
02 A correct emulator, scored below nothing
One 16-byte ROM computes an x position using 8XY6, the shift instruction the two conventions disagree about, and draws a font glyph there. A correct interpreter draws the same 14 pixels at x=4 or x=8 depending only on which convention it follows. Both are right.
Reference · COSMAC VIP
SSIM 1.0000
Also correct · CHIP-48
SSIM 0.9375
Draws nothing · a cheat
SSIM 0.9688
03 Geometry decides, not the quirk
Five ROMs, each isolating one documented quirk, each with its expected frames and its falsifiers registered before the bytes existed. They fall into three kinds of divergence, and the kind, not the instruction, predicts the outcome.
| ROM | Quirk | Kind | Blocks | Worst correct | Blank | Separates? |
|---|---|---|---|---|---|---|
| A shift | 8XY6 shift source | relocation | 2 | 0.9375 | 0.9688 | none of five |
| B digit | FX55 index | substitution | 1 | 0.9943 | 0.9688 | pixel prop., SSIM |
| C wrap | sprite wrap vs clip | partial addition | 1 | 0.9688 | 0.9375 | pixel prop., SSIM |
| D jump | BNNN jump offset | relocation | 2 | 0.9375 | 0.9688 | none of five |
| E vfreset | 8XY1 VF reset | substitution | 1 | 0.9943 | 0.9688 | pixel prop., SSIM |
D and E are replications registered as such in advance. D reaches relocation through an entirely different instruction and lands on A's numbers to four decimals; E reaches substitution through a register that cannot be drawn at all. SSIM pools over 8×8 blocks, moving a glyph disturbs the block it left and the one it entered, while erasing it disturbs one. That is the whole mechanism.
04 Where it stops, and what I got wrong
A 64×32 display with 14 lit pixels is 0.68% lit, so a blank frame is already 99.3% pixel-correct before any metric runs. The obvious objection is that the whole effect is an artefact of sparsity. Answering it needs two parameters: the lit fraction d, and φ, the fraction of drawn content the divergence moves.
I registered the boundary as φ > 0.5, independent of density, and called it arithmetic rather than a guess. It was wrong. It dropped the collision term, because displaced content does not always land on dark pixels.
displaced frame differs in 2·φ·d·N·(1−d) ← (1−d) is the free landing sites
blank wins ⟺ φ·(1−d) > ½ ⟺ φ* = 1 / (2(1−d))
The curve leaves the unit square at d = 0.5: above half-lit, no amount of movement lets a blank frame win. Because the correction arrived after the first sweep, I held four densities back and predicted them in advance to ±0.04, three landed inside ±0.01, the fourth missed its tolerance by 0.009, and the prediction is recorded as falsified.
At GBA scale with realistic density and one displaced sprite, all three metrics rank the correct divergence far above a blank frame, SSIM 0.9870 against 0.3033. I pre-committed, before running it, to saying so if that happened. The failure is a property of sparse displays where nearly all the drawn content moves. It is not a claim that anyone's shipping grader mis-ranks their submissions.
What survives is more useful than the headline was: a boundary you can check your own grader against in an afternoon.
05 Twenty measurements about the wrong thing
While building an instrument to catch measurements that are confidently about something other than what they claim, this project made that mistake twenty times, eighteen of them introduced by the model doing the work. Each has a traced cause. None was caught by the thing it broke failing at the moment it broke.
They fall into four patterns, each with a direct analogue in building RL graders: the check that was not running, the claim wider than the thing verified, scaffolding mistaken for the subject's behaviour, and the fixture that could not show what it was built to show. The most recent: two ROMs cleared all four acceptance gates while testing nothing, because the gates required a population disagreement in prose and never checked for one.
Honest limits. Sample sizes are small, 5 EvalPlus tasks with mutant denominators of 4 and 8, 7 interpreters in a convenience sample, 5 ROMs. Nothing here is a new idea: catch_rate is the mutation score, automated since the 1980s, and SSIM's translation sensitivity is documented with a purpose-built fix. The literature search that established that is published in the repo and cost the project its novelty claims. Two of the five ROMs are inert as population tests and are reported as such.