Attack taxonomy and credits¶
SealGrade does not invent a new classification. It uses the one published with BenchJack and cross-references the patterns from the Berkeley RDI write-up, then turns each class into runnable, regression-tested attacks.
Sources¶
- BenchJack: "Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack", arXiv 2605.12673. It defines the eight flaw classes V1 to V8 below and audits ten agent benchmarks with an LLM-driven pipeline.
- Berkeley RDI: "We Scored 100% on AI Benchmarks Without Solving a Single Problem", https://rdi.berkeley.edu/blog/trustworthy-benchmarks/. It describes seven recurring patterns, quoted here by short name.
Our contribution is different in kind: a deterministic runner, a corpus of concrete exploit payloads, and a regression matrix that fails CI when a defence stops working. We complement BenchJack's auditor; we do not replace it.
The eight classes (paraphrased)¶
| Class | Meaning |
|---|---|
| V1 | Isolation failure. The submission and the evaluator share an environment or a process. |
| V2 | Answers shipped with the test. The reference solution or expected outputs are reachable from where the submission runs. |
| V3 | Remote code execution into the evaluator. The evaluator parses or executes data the submission controls. |
| V4 | LLM-judge prompt injection. Output aimed at an LLM judge steers its score. |
| V5 | Weak string matching. Scorers that pattern-match can be satisfied by keyword stuffing. |
| V6 | Evaluation-logic gaps. A scorer can be satisfied by an easier property than the intended one. |
| V7 | Trusting untrusted output. Signals such as test output or reports can be influenced by the submission. |
| V8 | Excessive permissions. Root, writable host paths, or unrestricted network inside the sandbox. |
The seven Berkeley patterns (short names)¶
- Isolate everything that scores from everything being scored.
- Never trust output from the code you are evaluating.
- Test the tests, not just the submissions.
- Make tolerances and baselines honest.
- Treat evaluation code as production code.
- Answer leakage.
- Monkey-patching and stack introspection.
How the corpus maps to them¶
The corpus has 30 attacks (21 exploits and 9 boundary probes). The full list with each attack's classes and documented outcome per tier is generated from the manifests in CORPUS.md. In short:
| Class | Covered by |
|---|---|
| V1 isolation failure | tamper-tests-in-place, lingering-reward-writer, shadow-stdlib-argparse, pth-startup-hook, trojan-python-binary, frame-introspection-expected |
| V2 answers shipped with the test | read-answer-key-hardcode, runtime-read-answer-key |
| V3 execution into the evaluator | conftest-*, *-collect-only, atexit-forge-report, monkeypatch-pytest-reports |
| V4 LLM-judge injection | not exercised by the runner (it has no LLM judge); the auditor flags it (SG020) |
| V5 weak string matching | not a runner property; the auditor flags substring checks and loose tolerances (SG015, SG016) |
| V6 evaluation-logic gaps | always-equal-object, conftest-skip-all, pytest-ini-collect-only, nan-output |
| V7 trusting untrusted output | exit-zero-at-import, conftest-rewrite-junit, atexit-forge-report |
| V8 excessive permissions | trojan-python-binary, tamper-tests-in-place, pth-startup-hook, the artifact probes |
V4 and V5 are properties of how a task scores, not of the harness that runs it, so they are covered by auditor rules (and by the mutation score) rather than by runner attacks.
Using this responsibly¶
The payloads are written as test fixtures against SealGrade's own harness tiers. Do not point them at systems you do not own or have permission to test. See the security policy.