Auditor rules¶
sealgrade audit <path> runs these checks. They are heuristics over text and syntax
trees: specific (each finding names a file and line) and conservative (they prefer a low-severity note to a confident false alarm), but they can be wrong. SG027 and SG028 are produced by running things (--mutation, --dynamic) rather than by reading text.
| Rule | Severity | Classes | What it finds | How to fix |
|---|---|---|---|---|
SG001 |
high | V2 | Tests, solutions or answers are copied into the agent image. COPY/ADD places test or answer material in the image the agent runs in. | Keep tests and ground truth out of the agent image; mount them only into the verifier. |
SG002 |
medium | V2 | An answer-like file sits in the agent-visible directory. Files named expected/golden/answer* in the build context may be reachable by the agent. | Move ground truth out of the build context. |
SG003 |
medium | V8 | The agent runs as root. No non-root user for the agent. | Add a USER instruction or set [agent] user. |
SG004 |
high | V8 | A credential is baked into the environment. Secrets in the Dockerfile or .env files are readable by the agent. | Remove secrets; inject only at verification time into the verifier. |
SG005 |
medium | V3, V7 | The reward is derived from a process exit code alone. Exit status can be forced to 0. | Read a structured report; require the expected test count. |
SG006 |
high | V3, V7 | pytest configuration and conftest discovery are not pinned. pytest walks ancestors for config and conftest files. | Use -c, --rootdir and --confcutdir pointing at trusted files. |
SG007 |
medium | V3 | pytest plugin autoloading is enabled. Entry-point plugins load automatically. | Set PYTEST_DISABLE_PLUGIN_AUTOLOAD=1. |
SG008 |
medium | V1, V3 | The interpreter starts in an agent-writable directory. python -m puts the current directory first on sys.path. | Use python -I or cd elsewhere. |
SG009 |
medium | V7 | World-writable permissions on verifier paths. chmod 777 on /logs, /tests or /solution. | Keep verifier paths root-owned and read-only. |
SG010 |
medium | V7 | A test report is trusted without checking the test count. A zero-failure report is accepted regardless of how many tests it covers. | Require the exact expected count and no skips. |
SG011 |
high | V8 | The environment grants privileged or host access. privileged, docker.sock, host networking or added capabilities. | Remove them; they defeat container isolation. |
SG012 |
medium | V6 | The base image is floating or unpinned. latest tags or tags without digests. | Pin with @sha256 digests. |
SG013 |
medium | V6 | Packages are installed unpinned at verification time. Unpinned pip/npm installs in verifier scripts. | Pin exact versions or bake into the image. |
SG014 |
medium | V3, V7 | Remote code is piped into a shell. curl | sh at verification or build time. |
SG015 |
medium | V5 | Assertions check substrings instead of exact values. assert 'x' in output style checks. | Compare complete normalised values. |
SG016 |
medium | V5, V6 | A numeric tolerance is loose. approx/isclose tolerances of 1% or more, or absolute tolerance of 1. | Justify and tighten every tolerance. |
SG017 |
medium | V6 | Tests only check existence or type. All assertions in a test are existence/type/non-empty checks. | Assert on actual values. |
SG018 |
high | V6 | No effective verifier. No tests, or tests without assertions. | Add tests that compare outcomes. |
SG019 |
medium | V2 | .git history is agent-visible. A .git directory in the build context exposes history. | Exclude .git via .dockerignore. |
SG020 |
medium | V4 | An LLM judge is used or its prompt is built from agent output. LLM judges are injectable and non-deterministic. | Prefer deterministic checks; quote and isolate input. |
SG021 |
medium | V7 | Program output decides the reward. grep on program output sets the reward. | Use structured data from trusted code. |
SG022 |
low | V6 | Ground truth may be non-deterministic. Unseeded randomness or the clock feeds expected values. | Seed and fix inputs. |
SG023 |
low | V1, V7 | Candidate code is imported into the test process. Tests and candidate share a process. | Compare outputs of a separate process. |
SG024 |
medium | V1 | The agent and the verifier share one environment. Agent-phase changes persist into verification. | Use a separate verifier environment. |
SG025 |
low | V8 | The agent may have internet access. Network is not disabled. | Set network_mode = "none" unless required. |
SG026 |
low | V6 | Very few graded cases. Fewer than 20 cases. | Add edge cases and invalid input. |
SG027 |
medium | V6 | Weak verifier: low mutation score. Many plausible bugs in the reference solution still pass the tests. | Add cases that distinguish the surviving mutants. |
SG028 |
high | V1, V2, V3, V6, V7, V8 | An exploit from the corpus succeeds. A payload that does not solve the task obtained a passing verdict. | Move to a stricter tier or remove the weakness the attack targets. |