Skip to content

Auditor rules

sealgrade audit <path> runs these checks. They are heuristics over text and syntax trees: specific (each finding names a file and line) and conservative (they prefer a low-severity note to a confident false alarm), but they can be wrong. SG027 and SG028 are produced by running things (--mutation, --dynamic) rather than by reading text.

Rule Severity Classes What it finds How to fix
SG001 high V2 Tests, solutions or answers are copied into the agent image. COPY/ADD places test or answer material in the image the agent runs in. Keep tests and ground truth out of the agent image; mount them only into the verifier.
SG002 medium V2 An answer-like file sits in the agent-visible directory. Files named expected/golden/answer* in the build context may be reachable by the agent. Move ground truth out of the build context.
SG003 medium V8 The agent runs as root. No non-root user for the agent. Add a USER instruction or set [agent] user.
SG004 high V8 A credential is baked into the environment. Secrets in the Dockerfile or .env files are readable by the agent. Remove secrets; inject only at verification time into the verifier.
SG005 medium V3, V7 The reward is derived from a process exit code alone. Exit status can be forced to 0. Read a structured report; require the expected test count.
SG006 high V3, V7 pytest configuration and conftest discovery are not pinned. pytest walks ancestors for config and conftest files. Use -c, --rootdir and --confcutdir pointing at trusted files.
SG007 medium V3 pytest plugin autoloading is enabled. Entry-point plugins load automatically. Set PYTEST_DISABLE_PLUGIN_AUTOLOAD=1.
SG008 medium V1, V3 The interpreter starts in an agent-writable directory. python -m puts the current directory first on sys.path. Use python -I or cd elsewhere.
SG009 medium V7 World-writable permissions on verifier paths. chmod 777 on /logs, /tests or /solution. Keep verifier paths root-owned and read-only.
SG010 medium V7 A test report is trusted without checking the test count. A zero-failure report is accepted regardless of how many tests it covers. Require the exact expected count and no skips.
SG011 high V8 The environment grants privileged or host access. privileged, docker.sock, host networking or added capabilities. Remove them; they defeat container isolation.
SG012 medium V6 The base image is floating or unpinned. latest tags or tags without digests. Pin with @sha256 digests.
SG013 medium V6 Packages are installed unpinned at verification time. Unpinned pip/npm installs in verifier scripts. Pin exact versions or bake into the image.
SG014 medium V3, V7 Remote code is piped into a shell. curl sh at verification or build time.
SG015 medium V5 Assertions check substrings instead of exact values. assert 'x' in output style checks. Compare complete normalised values.
SG016 medium V5, V6 A numeric tolerance is loose. approx/isclose tolerances of 1% or more, or absolute tolerance of 1. Justify and tighten every tolerance.
SG017 medium V6 Tests only check existence or type. All assertions in a test are existence/type/non-empty checks. Assert on actual values.
SG018 high V6 No effective verifier. No tests, or tests without assertions. Add tests that compare outcomes.
SG019 medium V2 .git history is agent-visible. A .git directory in the build context exposes history. Exclude .git via .dockerignore.
SG020 medium V4 An LLM judge is used or its prompt is built from agent output. LLM judges are injectable and non-deterministic. Prefer deterministic checks; quote and isolate input.
SG021 medium V7 Program output decides the reward. grep on program output sets the reward. Use structured data from trusted code.
SG022 low V6 Ground truth may be non-deterministic. Unseeded randomness or the clock feeds expected values. Seed and fix inputs.
SG023 low V1, V7 Candidate code is imported into the test process. Tests and candidate share a process. Compare outputs of a separate process.
SG024 medium V1 The agent and the verifier share one environment. Agent-phase changes persist into verification. Use a separate verifier environment.
SG025 low V8 The agent may have internet access. Network is not disabled. Set network_mode = "none" unless required.
SG026 low V6 Very few graded cases. Fewer than 20 cases. Add edge cases and invalid input.
SG027 medium V6 Weak verifier: low mutation score. Many plausible bugs in the reference solution still pass the tests. Add cases that distinguish the surviving mutants.
SG028 high V1, V2, V3, V6, V7, V8 An exploit from the corpus succeeds. A payload that does not solve the task obtained a passing verdict. Move to a stricter tier or remove the weakness the attack targets.