Your grader, run as-is
Every candidate runs in a fresh container from the task’s image, with the network off and all capabilities dropped, using the task’s own eval script and grading.
A coding environment rewards a patch when its tests pass. When the tests are weak, a model learns to pass them instead of solving the task. Plumbline Grader finds the tasks that pay full reward for wrong code, and writes the tests that close them.
From our audit of a 200-task sample of SWE-bench Verified, the benchmark much of the field reports against. In each case the task’s own tests passed a patch that fails what the issue asks for. Count every wrong patch, not only the ones the issue rules out, and it is 143 of 191.
Each task runs through its own grader, unchanged, in a sealed container. We hand it wrong solutions and record which ones it pays for. Nothing counts as a finding until an independent check shows the patch is wrong.
Every candidate runs in a fresh container from the task’s image, with the network off and all capabilities dropped, using the task’s own eval script and grading.
Pieces of the reference fix with parts removed, needing no model, and plausible bugs and partial fixes written by a model. Each source is reported on its own.
A paid-out patch must differ from the reference fix in code. Then a separate check must pass on the reference fix, fail on the unfixed code, and fail on the patch.
Each finding is labeled: the issue as written requires the missing behavior (strict), or it is a real test gap outside the issue (broad). Both are reported, separately.
A sample of findings is re-run in fresh containers, check scripts included. Every number in the report traces to a run log, or the report is held.
New tests that pass the reference fix and fail every wrong patch we found, verified in the container, then attacked again with fresh wrong patches.
KEEP, FIX or DROP, with the wrong patch that passed, the hash of its run log, and the check that shows it is wrong.
A test patch for every FIX task, run against the reference fix and every known wrong patch before you see it.
We go through the findings with your team, task by task, and show how to reproduce each one.
We ran the auditor on public tasks first, so every claim here can be reproduced from the run artifacts.
For comparison, a June 2026 audit (Rajan) reported at least 28.5% on its 50-task sample of the same benchmark. On those same tasks our strict count is 17 of 48; theirs was 14 of 49.
A reward signal is only as good as the tests behind it. A model trained on a grader that pays for wrong code learns the wrong lesson, and the benchmark scores built on that grader stop meaning what they say.
A plumb line checks against a true vertical. We check every grader against the reference fix, not the tests’ own say-so.
Send us a sample of your environment’s tasks. We’ll show you what passes that shouldn’t, and how to close it.
Request an audit or write to shane@plumblinegrader.com