SealGrade¶
A tamper-resistant evaluation runner, an exploit corpus and an auditor for untrusted, AI-generated code.
Evaluation harnesses for AI agents usually trust the thing they are measuring. A submission that can write a
conftest.py, shadow a module, edit the tests, read the answer key or leave a process behind can score full
marks without solving anything. Public research has shown this on well-known benchmarks
(BenchJack, arXiv 2605.12673).
SealGrade is a deterministic answer:
- A runner with four harness tiers, from naive to strict, so the same attacks can be replayed against each.
- An exploit corpus of 30 labelled attacks. Each is a submission that does not solve the task, so a passing verdict means the exploit worked.
- A proof matrix that CI regenerates, so a defence regression fails the build.
- An auditor (
sealgrade audit) that scans task directories for the same flaw classes: 26 static rules, optional mutation scoring of the verifier, SARIF output for GitHub code scanning, and an adapter for Harbor / Terminal-Bench style tasks.
Where to start¶
| If you want to... | Read |
|---|---|
| See what it proves | Proof matrix |
| Understand the guarantees and their limits | Threat model |
| Harden your own evaluation | How to write an unhackable eval |
| Audit a task directory | Auditor rules |
| Add an attack | Exploit corpus and the contributing guide |
| See why it is built this way | ADR 0001 |
The four tiers¶
| Tier | How it grades | What it is for |
|---|---|---|
| T0 naive | one container; tests inside; reward from the exit code | the baseline every attack must beat |
| T1 typical | separate verifier container, same image, root, shared volumes, junit check | a common "better" setup that is still weak |
| T2 compat | isolated agent phase, artifact firewall, hardened pytest in a non-root, offline, read-only container | tasks whose tests must import the candidate |
| T3 strict | three containers; the judge compares data and never runs candidate code; signed verdicts | the recommended design |
T0 and T1 are our own re-creations of common patterns. They are not claims about any named benchmark or product.